You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python统计数据集文本列每行对应的全列词频

Python实现方案

我们可以分两步完成:先全局统计所有单词的总出现次数,再逐行匹配每个单词的总频率输出结果,这里提供常用的pandas处理方案,适合表格类数据集:

完整代码

import pandas as pd
from collections import Counter

# 1. 读取数据,如果是本地csv文件,替换成 pd.read_csv("你的文件路径.csv")即可
df = pd.DataFrame({
    "Text": [
        "This is a long string of words",
        "words have many types",
        "each type represents one thing",
        "thing are different",
        "where are these words"
    ]
})

# 2. 统计全列所有单词的总出现频率(默认不区分大小写)
all_words = []
for text in df["Text"]:
    # 拆分单词+转小写,需要区分大小写可去掉.lower()
    words = [word.lower() for word in text.split()]
    all_words.extend(words)
word_total_count = Counter(all_words)

# 3. 逐行生成每个单词的频率统计结果
def get_row_word_count(row_text):
    row_words = [w.lower() for w in row_text.split()]
    count_list = [f"{w}:{word_total_count[w]}" for w in row_words]
    return ", ".join(count_list)

df["Count"] = df["Text"].apply(get_row_word_count)

# 4. 输出/保存结果,保存到本地可替换为 df.to_csv("结果文件.csv", index=False)
print(df)

输出结果示例

运行上述代码后输出的结果格式如下:

Text                                                                 Count
0   This is a long string of words  this:1, is:1, a:1, long:1, string:1, of:1, words:3
1            words have many types                           words:3, have:1, many:1, types:1
2  each type represents one thing     each:1, type:1, represents:1, one:1, thing:2
3              thing are different                                          thing:2, are:2, different:1
4            where are these words                               where:1, are:2, these:1, words:3

自定义调整说明

  • 如果需要区分单词大小写,删除代码中所有.lower()即可
  • 如果文本里有标点符号需要过滤,可以在拆分单词时增加正则匹配替换,比如用re.sub(r'[^\w\s]', '', text)先去掉标点再拆分
  • 结果可以直接调用df.to_csv()导出为本地csv文件使用

内容的提问来源于stack exchange,提问作者Kath

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 18:00:01