使用collections与pandas DataFrame统计唯一词仅统计首句的问题
问题描述
统计pandas DataFrame文本列的所有唯一词时,代码运行后仅统计第一行文本内容,未遍历处理全部文本数据。
测试使用的DataFrame数据如下:
text 0 hello is a unique sentences 1 hello this is a test 2 does this works
数据构造代码:
import pandas as pd d = { "text": ["hello is a unique sentences", "hello this is a test", "does this works"], } df = pd.DataFrame(data=d)
原有实现代码:
from collections import Counter # Count unique words def counter_word(text_col): print(len(text_col.values)) count = Counter() for i, text in enumerate(text_col.values): print(i) for word in text.split(): count[word] += 1 return count counter = counter_word(df['text']) len(counter)
错误原因
核心问题是return count语句缩进错误:该语句被放在了遍历文本列的外层for循环内部,函数处理完索引为0的第一行文本后,就会直接执行return终止函数运行,后续两行文本完全不会进入遍历流程,因此最终只统计到第一行的单词。
修复方案
将return count的缩进调整至和外层for循环同级,等所有行的文本都遍历统计完成后再返回计数结果。
修复后的完整代码:
import pandas as pd from collections import Counter d = { "text": ["hello is a unique sentences", "hello this is a test", "does this works"], } df = pd.DataFrame(data=d) def counter_word(text_col): count = Counter() for text in text_col.values: for word in text.split(): count[word] += 1 return count # 移到循环外部,全量遍历完成后再返回 counter = counter_word(df['text']) print(len(counter)) # 输出结果为9,即全量文本去重后共9个唯一词
简化实现
如果不需要逐行打印调试信息,可以直接拼接所有文本后一次性统计,代码更简洁:
from collections import Counter counter = Counter(" ".join(df['text']).split())
内容的提问来源于stack exchange,提问作者Test
相关产品推荐
相关产品推荐

