处理百万级Twitter XML数据集时触发AttributeError: 'NoneType' object has no attribute 'count'问题排查
问题诊断与解决方案:全量XML数据集处理时的NoneType错误
问题背景
你正在处理包含100万条推文的XML数据集,目标是提取推文中的引号内容并生成CSV文件。测试5万条样本数据时代码运行正常,但处理全量数据时触发了以下错误:
Traceback (most recent call last): File "C:/xxx.py", line 16, in
count = text.count("'") AttributeError: 'NoneType' object has no attribute 'count'
错误原因
这个错误的核心原因很明确:你的全量XML数据中存在空的<paragraph>标签,而5万条样本数据刚好没有这种情况。
当XML里出现<paragraph></paragraph>或者自闭合的<paragraph/>标签时,paragraph.text会返回None(而不是空字符串)。你的代码直接对text调用count()方法,当text是None时自然会触发AttributeError。
解决方案
只需要在处理text之前增加一个空值判断,跳过空段落或者将None转为空字符串即可。下面是修改后的代码,重点标注了新增的判断逻辑:
import xml.etree.ElementTree as ET import pandas as pd import numpy as np tree = ET.parse('tweets.xml') articles = tree.getroot() paragraphs_with_quotes = [] paragraphs_with_double_quotes = [] quotes = [] extracted_paragraphs = [] for article in articles: for paragraph in article.findall('paragraph'): text = paragraph.text # 新增:处理text为None的情况 if text is None: continue # 直接跳过空段落 # 或者如果你想保留空段落记录,可以改为:text = text or "" count = text.count("\'") indexes = [] if count > 1: paragraphs_with_quotes.append(text) index = text.index("\'") while count > 0: if text[index - 1] == " " or index == len(text) - 1 or text[index + 1] in " .,:": indexes.append(index) if count > 1: index = text.index("\'", index + 1) count -= 1 for i in range(0, len(indexes), 2): start = indexes[i] end = indexes[min(len(indexes) - 1, i + 1)] print(text) quotes.append(text[indexes[i]:indexes[min(len(indexes) - 1, i + 1)] + 1]) extracted_paragraphs.append(text) print("Quote:" + quotes[len(quotes) - 1]) print() d = {'Paragraph:': extracted_paragraphs, 'Quote:': quotes} quote_data = pd.DataFrame(d) quote_data.to_csv('quote_data.csv') for i in range(1): print() print(len(paragraphs_with_quotes))
额外建议
考虑到你的数据集有100万条记录,还可以做以下优化提升效率:
- 避免在循环中频繁调用
print(),这会大幅拖慢处理速度,建议只在调试时保留 - 使用生成器来逐步处理数据,减少内存占用
- 对引号提取逻辑可以用正则表达式简化,比如用
re.findall(r"'([^']+)'", text)来匹配单引号包裹的内容,代码会更简洁
内容的提问来源于stack exchange,提问作者klassetyp
相关产品推荐
相关产品推荐

