You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量处理XML文档分词后结果顺序不符的问题排查

XML文件解析后分词列表与预期顺序不符的问题

我用以下代码解析1400个XML文件,生成每个文件的去停用词、词干化分词列表:

content = []

path = 'assets/cranfield/cranfieldDocs/'
files = os.listdir(path)
stop_list = stopwords.words('english')
st = PorterStemmer()
for file in files:
    tree = ET.parse(os.path.join(path + file))
    root = tree.getroot()
    for idx,log_element in enumerate(root.findall('TEXT')):
        tokens = word_tokenize(log_element.text)
        clean_tokens = [word for word in tokens if word not in stop_list]
        stem_tokens = [st.stem(word) for word in clean_tokens]
        content.append(stem_tokens)

但输出结果不符合预期:
示例第一个文档(DOCNO为1)的TEXT内容如下:

<DOC>
<DOCNO>
1
</DOCNO>
<TITLE>
experimental investigation of the aerodynamics of a
wing in a slipstream .
</TITLE>
<AUTHOR>
brenckman,m.
</AUTHOR>
<BIBLIO>
j. ae. scs. 25, 1958, 324.
</BIBLIO>
<TEXT>
  an experimental study of a wing in a propeller slipstream was
made in order to determine the spanwise distribution of the lift
increase due to slipstream at different angles of attack of the wing
and at different free stream to slipstream velocity ratios .  the
results were intended in part as an evaluation basis for different
theoretical treatments of this problem .
  the comparative span loading curves, together with supporting
evidence, showed that a substantial part of the lift increment
produced by the slipstream was due to a /destalling/ or boundary-layer-control
effect .  the integrated remaining lift increment,
after subtracting this destalling lift, was found to agree
well with a potential flow theory .
  an empirical evaluation of the destalling effects was made for
the specific configuration of the experiment .
</TEXT>
</DOC>

预期第一个分词列表为:['experiment', 'studi', ..., 'configur', 'experi', '.']
但实际得到的是:['reciproc', 'equat', ..., 'might', 'conserv', '.']


原因分析

  • os.listdir()返回的文件列表不遵循固定的排序规则,其顺序由操作系统的文件系统决定,并非按文件名的数字顺序或DOCNO顺序排列。你代码中遍历的第一个文件并不是DOCNO为1的目标文件,因此content列表的第一个元素对应其他文档的分词结果。

解决方案

  • 对文件列表按文件名的数字部分排序(假设文件名为1.xml、2.xml这类格式):
files = sorted(os.listdir(path), key=lambda x: int(os.path.splitext(x)[0]))
  • 若文件名格式不同,可根据实际情况调整排序逻辑,确保遍历顺序与DOCNO的顺序一致。

内容的提问来源于stack exchange,提问作者Hefe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 17:57:25