You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

处理百万级Twitter XML数据集时触发AttributeError: 'NoneType' object has no attribute 'count'问题排查

问题诊断与解决方案:全量XML数据集处理时的NoneType错误

问题背景

你正在处理包含100万条推文的XML数据集,目标是提取推文中的引号内容并生成CSV文件。测试5万条样本数据时代码运行正常,但处理全量数据时触发了以下错误:

Traceback (most recent call last): File "C:/xxx.py", line 16, in count = text.count("'") AttributeError: 'NoneType' object has no attribute 'count'

错误原因

这个错误的核心原因很明确:你的全量XML数据中存在空的<paragraph>标签,而5万条样本数据刚好没有这种情况。

当XML里出现<paragraph></paragraph>或者自闭合的<paragraph/>标签时,paragraph.text会返回None(而不是空字符串)。你的代码直接对text调用count()方法,当text是None时自然会触发AttributeError。

解决方案

只需要在处理text之前增加一个空值判断,跳过空段落或者将None转为空字符串即可。下面是修改后的代码,重点标注了新增的判断逻辑:

import xml.etree.ElementTree as ET
import pandas as pd
import numpy as np

tree = ET.parse('tweets.xml')
articles = tree.getroot()

paragraphs_with_quotes = []
paragraphs_with_double_quotes = []
quotes = []
extracted_paragraphs = []

for article in articles:
    for paragraph in article.findall('paragraph'):
        text = paragraph.text
        # 新增:处理text为None的情况
        if text is None:
            continue  # 直接跳过空段落
        # 或者如果你想保留空段落记录,可以改为:text = text or ""
        
        count = text.count("\'")
        indexes = []
        if count > 1:
            paragraphs_with_quotes.append(text)
            index = text.index("\'")
            while count > 0:
                if text[index - 1] == " " or index == len(text) - 1 or text[index + 1] in " .,:":
                    indexes.append(index)
                if count > 1:
                    index = text.index("\'", index + 1)
                count -= 1
            for i in range(0, len(indexes), 2):
                start = indexes[i]
                end = indexes[min(len(indexes) - 1, i + 1)]
                print(text)
                quotes.append(text[indexes[i]:indexes[min(len(indexes) - 1, i + 1)] + 1])
                extracted_paragraphs.append(text)
                print("Quote:" + quotes[len(quotes) - 1])
                print()

d = {'Paragraph:': extracted_paragraphs, 'Quote:': quotes}
quote_data = pd.DataFrame(d)
quote_data.to_csv('quote_data.csv')

for i in range(1):
    print()
print(len(paragraphs_with_quotes))

额外建议

考虑到你的数据集有100万条记录,还可以做以下优化提升效率:

  • 避免在循环中频繁调用print(),这会大幅拖慢处理速度,建议只在调试时保留
  • 使用生成器来逐步处理数据,减少内存占用
  • 对引号提取逻辑可以用正则表达式简化,比如用re.findall(r"'([^']+)'", text)来匹配单引号包裹的内容,代码会更简洁

内容的提问来源于stack exchange,提问作者klassetyp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 14:52:42