You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计列表单词在语料库行中的出现次数并修正代码

修正单词行计数代码 + 更简便实现方案

需求是统计列表c中每个单词在corpus各行的出现行数(只要该行包含该单词就算一次,不管出现几次),期望输出为[1,3,2,4,1,1,4,1,4]。

原代码问题分析

原代码的核心错误是在统计前就把初始值0添加到结果列表,导致统计后的正确count没有被存入,最终结果全是错位的无效值。比如处理第一个单词"and"时,先把0加入a,然后统计得到count=1,但这个1没被添加,结果列表的第一个元素就成了错误的0。

修正后的代码

把a.append(count)移到内层循环结束后,确保统计完成再把结果加入列表:

corpus= [
         'this is the first document',            #row1
         'this document is the second document',  #row2
         'and this is the third one',             #row3
         'is this the first document',            #row4
    ]

c=['and', 'document', 'first', 'is', 'one', 'second', 'the', 'third', 'this']

a=[]

for word in c:
    count = 0
    for row in corpus:
        if word in row:
            count += 1
    a.append(count)  # 统计完成后再添加结果

print(a)  # 输出:[1, 3, 2, 4, 1, 1, 4, 1, 4]

更简便的实现方案

一行式列表推导方案

利用Python的列表推导式+生成器表达式,逻辑简洁直接:

corpus= [
         'this is the first document',
         'this document is the second document',
         'and this is the third one',
         'is this the first document',
    ]

c=['and', 'document', 'first', 'is', 'one', 'second', 'the', 'third', 'this']

result = [sum(1 for row in corpus if word in row) for word in c]
print(result)  # 输出:[1, 3, 2, 4, 1, 1, 4, 1, 4]

原理:对c中的每个单词,遍历corpus的每一行,判断单词是否在该行中,存在则返回1,否则0,用sum()把这些1累加就是该单词出现的行数。

大语料优化方案

如果corpus规模很大,可以先把每行的单词转成集合,把查找时间复杂度从O(n)降到O(1),提升效率:

# 预处理corpus,把每行转成单词集合
corpus_sets = [set(row.split()) for row in corpus]
result = [sum(1 for s in corpus_sets if word in s) for word in c]

内容的提问来源于stack exchange,提问作者buzz bowlekar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 03:24:25