You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计CSV列中所有单词的频率并识别重叠词?

需求可行性分析与实现思路

完全可行,不需要预先指定特定单词就能实现CSV某一列所有单词的频率统计,同时识别单元格内的重叠单词,以下是具体实现思路:

核心实现步骤

  • 读取并提取目标列数据
    用Python的pandas库读取CSV文件,提取需要分析的列并过滤空值:

    import pandas as pd
    df = pd.read_csv("your_data.csv")
    target_column = df["目标列名称"].dropna().tolist()
    
  • 文本预处理
    统一文本格式、去除标点符号、拆分单词,处理特殊字符(比如示例中的Australia’s):

    import re
    from nltk.tokenize import word_tokenize
    
    processed_texts = []
    for text in target_column:
        # 转小写、去除标点
        cleaned = re.sub(r'[^\w\s\']', '', text.lower())
        # 分词
        tokens = word_tokenize(cleaned)
        processed_texts.extend(tokens)
    
  • 统计所有单词频率
    用collections.Counter直接统计所有单词的出现次数:

    from collections import Counter
    word_frequency = Counter(processed_texts)
    # 按频率从高到低输出
    for word, count in sorted(word_frequency.items(), key=lambda x: x[1], reverse=True):
        print(f"{word}: {count}")
    
  • 识别单元格内的重叠单词
    针对单个单元格内的连续重复单词(比如类似"aged care care homes"中的care),可以用正则匹配检测:

    # 正则匹配连续重复单词
    overlap_pattern = re.compile(r'(\b\w+\b)\s+\1')
    overlap_results = []
    for text in target_column:
        matches = overlap_pattern.findall(text.lower())
        if matches:
            overlap_results.append({"原文本": text, "重叠单词": matches})
    print(overlap_results)
    

示例文本翻译

  1. 澳大利亚养老院迎来严峻里程碑;
  2. 针对养老院的严厉审计揭露了虐待与忽视问题。

内容的提问来源于stack exchange,提问作者Mr. Student.exe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 17:31:04