You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用collections与pandas DataFrame统计唯一词仅统计首句的问题

问题描述

统计pandas DataFrame文本列的所有唯一词时,代码运行后仅统计第一行文本内容,未遍历处理全部文本数据。
测试使用的DataFrame数据如下:

text
0  hello is a unique sentences
1         hello this is a test
2              does this works

数据构造代码:

import pandas as pd
d = {
    "text": ["hello is a unique sentences",
             "hello this is a test", 
             "does this works"],
}
df = pd.DataFrame(data=d)

原有实现代码:

from collections import Counter

# Count unique words
def counter_word(text_col):
    print(len(text_col.values))
    count = Counter()
    for i, text in enumerate(text_col.values):
        print(i)
        for word in text.split():
            count[word] += 1
        return count

counter = counter_word(df['text'])
len(counter)
错误原因

核心问题是return count语句缩进错误:该语句被放在了遍历文本列的外层for循环内部,函数处理完索引为0的第一行文本后,就会直接执行return终止函数运行,后续两行文本完全不会进入遍历流程,因此最终只统计到第一行的单词。

修复方案

将return count的缩进调整至和外层for循环同级,等所有行的文本都遍历统计完成后再返回计数结果。
修复后的完整代码:

import pandas as pd
from collections import Counter

d = {
    "text": ["hello is a unique sentences",
             "hello this is a test", 
             "does this works"],
}
df = pd.DataFrame(data=d)

def counter_word(text_col):
    count = Counter()
    for text in text_col.values:
        for word in text.split():
            count[word] += 1
    return count  # 移到循环外部,全量遍历完成后再返回

counter = counter_word(df['text'])
print(len(counter)) 
# 输出结果为9,即全量文本去重后共9个唯一词
简化实现

如果不需要逐行打印调试信息,可以直接拼接所有文本后一次性统计,代码更简洁:

from collections import Counter
counter = Counter(" ".join(df['text']).split())

内容的提问来源于stack exchange,提问作者Test

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 11:15:18