如何使用Python统计数据集文本列每行对应的全列词频
Python实现方案
我们可以分两步完成:先全局统计所有单词的总出现次数,再逐行匹配每个单词的总频率输出结果,这里提供常用的pandas处理方案,适合表格类数据集:
完整代码
import pandas as pd from collections import Counter # 1. 读取数据,如果是本地csv文件,替换成 pd.read_csv("你的文件路径.csv")即可 df = pd.DataFrame({ "Text": [ "This is a long string of words", "words have many types", "each type represents one thing", "thing are different", "where are these words" ] }) # 2. 统计全列所有单词的总出现频率(默认不区分大小写) all_words = [] for text in df["Text"]: # 拆分单词+转小写,需要区分大小写可去掉.lower() words = [word.lower() for word in text.split()] all_words.extend(words) word_total_count = Counter(all_words) # 3. 逐行生成每个单词的频率统计结果 def get_row_word_count(row_text): row_words = [w.lower() for w in row_text.split()] count_list = [f"{w}:{word_total_count[w]}" for w in row_words] return ", ".join(count_list) df["Count"] = df["Text"].apply(get_row_word_count) # 4. 输出/保存结果,保存到本地可替换为 df.to_csv("结果文件.csv", index=False) print(df)
输出结果示例
运行上述代码后输出的结果格式如下:
Text Count 0 This is a long string of words this:1, is:1, a:1, long:1, string:1, of:1, words:3 1 words have many types words:3, have:1, many:1, types:1 2 each type represents one thing each:1, type:1, represents:1, one:1, thing:2 3 thing are different thing:2, are:2, different:1 4 where are these words where:1, are:2, these:1, words:3
自定义调整说明
- 如果需要区分单词大小写,删除代码中所有
.lower()即可 - 如果文本里有标点符号需要过滤,可以在拆分单词时增加正则匹配替换,比如用
re.sub(r'[^\w\s]', '', text)先去掉标点再拆分 - 结果可以直接调用
df.to_csv()导出为本地csv文件使用
内容的提问来源于stack exchange,提问作者Kath
相关产品推荐
相关产品推荐

