You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何使用While循环分析CSV文件数据完成统计需求

基于while循环的Reddit CSV数据统计实现

你已经通过csv.DictReader把所有帖子数据解析成了固定结构的entries列表,每个元素对应一条帖子的(id, 得分, 评论数, 标题)数据,接下来直接用while循环遍历这个列表完成所有指标统计即可,无需二次读取文件。

完整可运行代码

import csv

def analyze(entries):
    # 兼容空数据场景
    if not entries:
        print("无有效帖子数据")
        return 0
    
    # 初始化统计变量
    total_score = 0
    total_comms = 0
    entry_count = len(entries)
    # 初始化极值相关变量,默认取第一个条目值
    max_score = entries[0][1]
    max_score_title = entries[0][3]
    min_score = entries[0][1]
    min_score_title = entries[0][3]
    max_comms = entries[0][2]
    max_comms_title = entries[0][3]
    # while循环控制索引
    i = 0

    while i < entry_count:
        current_id, current_score, current_comms, current_title = entries[i]
        # 累计总数据用于计算平均值
        total_score += current_score
        total_comms += current_comms

        # 更新最高得分记录
        if current_score > max_score:
            max_score = current_score
            max_score_title = current_title
        # 更新最低得分记录
        if current_score < min_score:
            min_score = current_score
            min_score_title = current_title
        # 更新最多评论记录
        if current_comms > max_comms:
            max_comms = current_comms
            max_comms_title = current_title
        
        i += 1
    
    # 计算平均值
    avg_score = total_score / entry_count
    avg_comms = total_comms / entry_count

    # 输出所有要求的统计结果
    print("=====Reddit帖子统计结果=====")
    print(f"所有帖子平均评论数:{avg_comms:.2f}")
    print(f"所有帖子平均得分:{avg_score:.2f}")
    print(f"最高得分对应帖子标题:{max_score_title}(得分:{max_score})")
    print(f"最低得分对应帖子标题:{min_score_title}(得分:{min_score})")
    print(f"评论数最多的帖子标题:{max_comms_title}(评论数:{max_comms})")

    return avg_score

with open("reddit_vm.csv", "r", encoding='UTF-8', errors="ignore") as input:
    entries = [(e['id'], int(e['score']), int(e['comms_num']), e['title']) for e in csv.DictReader(input)]
    avgScore = analyze(entries)

逻辑说明

  • 累计类统计变量初始值设为0,极值类变量默认取第一条数据的值,避免初始值偏差
  • 用索引变量i控制while循环遍历所有条目,每次循环同步更新所有统计值
  • 平均值保留2位小数输出,可根据需求调整精度

内容的提问来源于stack exchange,提问作者Spencer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 18:48:03