Python按标签拆分大型XML文件遇AttributeError报错求助
问题分析与解决
你的错误AttributeError: 'NoneType' object has no attribute 'text'主要来自两个核心问题:
1. 节点路径错误
从XML结构能看到,talkid是嵌套在head标签内部的,并非file的直接子元素。你用elem.find('talkid')无法定位到该节点,会返回None,调用.text自然触发报错。
2. 文件写入逻辑错误
- 你已经通过
.text获取了content的文本内容,不需要再用ET.tostring()——这个方法是用来序列化XML元素的,不适合处理纯文本。 - 打开文件用
wb二进制模式时,不能指定encoding参数;如果要写入文本内容,应该使用w模式。
修正后的代码
import xml.etree.ElementTree as ET all_talks = 'path\\to\\big\\file' context = ET.iterparse(all_talks, events=('end', )) for event, elem in context: if elem.tag == 'file': # 先定位head节点,再找内部的talkid head_node = elem.find('head') if not head_node: elem.clear() continue talkid_node = head_node.find('talkid') if not talkid_node or not talkid_node.text: elem.clear() continue # 获取content节点的文本内容 content_node = elem.find('content') if not content_node or not content_node.text: elem.clear() continue title = talkid_node.text.strip() filename = f"{title}.txt" # 用文本模式写入并指定编码 with open(filename, 'w', encoding='utf-8') as f: f.write(content_node.text.strip()) # 清理元素释放内存,处理大XML文件必备操作 elem.clear()
额外说明
- 加入了空节点判断,避免部分
file节点缺失head、talkid或content时程序崩溃。 elem.clear()是处理大型XML的关键,能避免内存占用过高。- 用
strip()去除文本前后的空白字符,让输出内容更整洁。
内容的提问来源于stack exchange,提问作者Leila
相关产品推荐
相关产品推荐

