Python解析XML转pandas DataFrame报错及输出不符问题求助
问题原因&解决方法
第一个报错AttributeError: 'NoneType' object has no attribute 'find_all'的核心原因是:你的XML结构里没有名为RECORDING的标签,原有代码是适配根标签为RECORDING的XML文件,和你的实际数据结构不匹配。
替换成soup.find_all后仍不符合预期,是因为原有逻辑只提取了标签的文本内容,但你需要的category、polarity是aspectCategory标签的属性值,且一个句子对应多条情感分类结果,需要按分类拆分为多行。
调整后可直接运行的代码
import pandas as pd from bs4 import BeautifulSoup # 读取解析XML文件 with open("C:/Users/user/Desktop/essai/dataXml.xml", "r", encoding="utf-8") as f: soup = BeautifulSoup(f, "xml") res_list = [] # 遍历所有sentence标签 for sentence in soup.find_all("sentence"): # 提取句子文本 text = sentence.find("text").get_text(strip=True) # 遍历当前句子下的所有aspectCategory标签 for aspect in sentence.find_all("aspectCategory"): # 提取属性值拼接为单条记录 res_list.append({ "text": text, "category": aspect.get("category"), "polarity": aspect.get("polarity") }) # 转换为DataFrame df = pd.DataFrame(res_list) print(df)
输出效果
| text | category | polarity |
|---|---|---|
| We went again and sat at the bar this time, I had 5 pints of guinness and not one buy-back, I ordered a basket of onion rings and there were about 5 in the basket, the rest was filled with crumbs, the chili was not even edible. | place | neutral |
| We went again and sat at the bar this time, I had 5 pints of guinness and not one buy-back, I ordered a basket of onion rings and there were about 5 in the basket, the rest was filled with crumbs, the chili was not even edible. | food | negative |
内容的提问来源于stack exchange,提问作者mina
相关产品推荐
相关产品推荐

