如何统计列表单词在语料库行中的出现次数并修正代码
修正单词行计数代码 + 更简便实现方案
需求是统计列表c中每个单词在corpus各行的出现行数(只要该行包含该单词就算一次,不管出现几次),期望输出为[1,3,2,4,1,1,4,1,4]。
原代码问题分析
原代码的核心错误是在统计前就把初始值0添加到结果列表,导致统计后的正确count没有被存入,最终结果全是错位的无效值。比如处理第一个单词"and"时,先把0加入a,然后统计得到count=1,但这个1没被添加,结果列表的第一个元素就成了错误的0。
修正后的代码
把a.append(count)移到内层循环结束后,确保统计完成再把结果加入列表:
corpus= [ 'this is the first document', #row1 'this document is the second document', #row2 'and this is the third one', #row3 'is this the first document', #row4 ] c=['and', 'document', 'first', 'is', 'one', 'second', 'the', 'third', 'this'] a=[] for word in c: count = 0 for row in corpus: if word in row: count += 1 a.append(count) # 统计完成后再添加结果 print(a) # 输出:[1, 3, 2, 4, 1, 1, 4, 1, 4]
更简便的实现方案
一行式列表推导方案
利用Python的列表推导式+生成器表达式,逻辑简洁直接:
corpus= [ 'this is the first document', 'this document is the second document', 'and this is the third one', 'is this the first document', ] c=['and', 'document', 'first', 'is', 'one', 'second', 'the', 'third', 'this'] result = [sum(1 for row in corpus if word in row) for word in c] print(result) # 输出:[1, 3, 2, 4, 1, 1, 4, 1, 4]
原理:对c中的每个单词,遍历corpus的每一行,判断单词是否在该行中,存在则返回1,否则0,用sum()把这些1累加就是该单词出现的行数。
大语料优化方案
如果corpus规模很大,可以先把每行的单词转成集合,把查找时间复杂度从O(n)降到O(1),提升效率:
# 预处理corpus,把每行转成单词集合 corpus_sets = [set(row.split()) for row in corpus] result = [sum(1 for s in corpus_sets if word in s) for word in c]
内容的提问来源于stack exchange,提问作者buzz bowlekar
相关产品推荐
相关产品推荐

