You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取CSV文件指定行数内容并实现单词按词频排序

解决CSV指定行数的单词词频统计问题

嘿,我来帮你搞定这个CSV词频统计的问题!你之前的代码没达到预期效果,主要是因为没正确处理行读取和单词拆分的逻辑,咱们一步步来解决:

问题分析

  • 你直接用collections.Counter(y)时,y是文件对象,Counter会把整行内容当作单个元素统计,而不是拆分每行里的单词,这显然不是你要的结果。
  • 用CSV Reader的话,你需要明确逐行读取、提取每行的单词,再统一统计,而不是直接把Reader对象传给Counter。

解决方案代码

下面是完整的实现,能读取指定行数的内容,统计所有单词的词频并按频率降序排列:

from collections import Counter
import csv

def calculate_word_frequency(filename, total_lines):
    # 存储所有要统计的单词
    collected_words = []
    
    # 安全打开文件,自动处理资源关闭
    with open(filename, 'r', newline='', encoding='utf-8') as file:
        csv_reader = csv.reader(file)
        
        # 读取指定行数,超出后停止
        for line_num, row in enumerate(csv_reader):
            if line_num >= total_lines:
                break
            # 过滤掉空字符串(可选,根据你的CSV情况调整)
            valid_words = [word for word in row if word.strip()]
            collected_words.extend(valid_words)
    
    # 统计词频并按频率从高到低排序
    word_counter = Counter(collected_words)
    sorted_word_frequencies = sorted(word_counter.items(), key=lambda item: item[1], reverse=True)
    
    return sorted_word_frequencies

# 使用示例:读取前30行内容并统计
if __name__ == "__main__":
    result = calculate_word_frequency("your_csv_file.csv", 30)
    print(result)

代码说明

  1. 文件读取与行控制:用csv.reader解析每行,通过enumerate跟踪行数,读取到total_lines后停止循环。
  2. 单词收集:把每行拆分后的单词(过滤掉空字符串)加入总列表,确保所有指定行的单词都被统计。
  3. 词频统计与排序:用Counter快速统计词频,再通过sorted按词频倒序排列,得到你需要的[('word1', freq1), ...]格式。

可选调整

如果你不需要过滤空字符串,直接把valid_words = ...那行换成collected_words.extend(row)即可。

内容的提问来源于stack exchange,提问作者SVP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 21:24:06