You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现CSV冗余词合并、去停用词并按作者统计词频

Python实现CSV词频统计(合并重复词、按作者拆分、去停用词)

需求梳理

你需要处理的是一个包含单词、作者、词频的空格分隔类CSV文件,核心要求如下:

  • 合并同一单词的所有记录,按作者累计词频
  • 移除英文停用词(比如OF、AND、IN这类高频无意义词汇)
  • 最终输出按作者拆分的词频表格,每行是单词,列是各作者的累计词频(无记录则填0)

完整实现代码

from collections import defaultdict

# 自定义英文停用词列表(覆盖你示例中的停用词,可根据需要补充)
STOPWORDS = {
    'OF', 'AND', 'IN', 'THE', 'A', 'TO', 'WITH', 'FOR', 'BY', 'AN', '1'
}

def process_word_frequency(input_file_path):
    # 初始化统计字典:word -> {author: total_freq}
    word_author_freq = defaultdict(lambda: defaultdict(int))
    all_authors = set()

    # 读取并处理每行数据
    with open(input_file_path, 'r', encoding='utf-8') as f:
        for line in f:
            line = line.strip()
            if not line:
                continue
            parts = line.split()
            # 处理作者名可能包含空格的情况:第一个元素是单词,最后一个是词频,中间是作者名
            word = parts[0].upper()  # 统一大写,避免大小写差异导致的重复
            freq = int(parts[-1])
            author = ' '.join(parts[1:-1])
            all_authors.add(author)

            # 跳过停用词
            if word in STOPWORDS:
                continue

            # 累加词频
            word_author_freq[word][author] += freq

    # 将作者集合转为排序后的列表(保证输出顺序一致)
    sorted_authors = sorted(all_authors)

    # 构建输出内容
    output_lines = []
    # 表头:Words + 所有作者
    header = ['Words'] + sorted_authors
    output_lines.append('\t'.join(header))

    # 遍历每个单词,生成对应行
    for word in sorted(word_author_freq.keys()):
        row = [word]
        for author in sorted_authors:
            row.append(str(word_author_freq[word].get(author, 0)))
        output_lines.append('\t'.join(row))

    return '\n'.join(output_lines)

# 示例使用
if __name__ == '__main__':
    # 替换为你的输入文件路径
    input_path = 'your_input_file.txt'  # 因为是空格分隔,也可以用txt后缀
    result = process_word_frequency(input_path)
    print(result)
    # 可选:保存结果到文件
    with open('word_frequency_result.txt', 'w', encoding='utf-8') as f:
        f.write(result)

代码细节解释

  1. 停用词处理:

    • 自定义了停用词集合,包含你示例中出现的所有无意义词汇,你可以根据需求随时添加或删除。
    • 统一将单词转为大写,避免THE和the被视为不同单词的情况。
  2. 作者名处理:

    • 考虑到作者名可能包含空格(比如Hamzad Ali),代码通过' '.join(parts[1:-1])拼接中间部分作为完整作者名,避免拆分错误。
  3. 统计逻辑:

    • 使用嵌套的defaultdict,外层键是单词,内层键是作者,值是累计词频,方便自动初始化和累加。
    • 收集所有出现过的作者,确保输出表头包含所有作者。
  4. 输出格式:

    • 用制表符\t分隔列,保证输出的表格对齐,也方便后续导入Excel等工具处理。
    • 单词和作者都按排序输出,结果更规整。

输出示例(简化版)

Words	Pandey P	Hamzad Ali	Karen Sara	John christopher
HIV-1	4	49	0	0
HEPATITIS	0	13	34	13
INHIBITORS	0	39	19	0
KINASE	31	0	0	0
HCV	0	0	23	0

内容的提问来源于stack exchange,提问作者spideypack

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:39:05