Python中如何使用While循环分析CSV文件数据完成统计需求
基于while循环的Reddit CSV数据统计实现
你已经通过csv.DictReader把所有帖子数据解析成了固定结构的entries列表,每个元素对应一条帖子的(id, 得分, 评论数, 标题)数据,接下来直接用while循环遍历这个列表完成所有指标统计即可,无需二次读取文件。
完整可运行代码
import csv def analyze(entries): # 兼容空数据场景 if not entries: print("无有效帖子数据") return 0 # 初始化统计变量 total_score = 0 total_comms = 0 entry_count = len(entries) # 初始化极值相关变量,默认取第一个条目值 max_score = entries[0][1] max_score_title = entries[0][3] min_score = entries[0][1] min_score_title = entries[0][3] max_comms = entries[0][2] max_comms_title = entries[0][3] # while循环控制索引 i = 0 while i < entry_count: current_id, current_score, current_comms, current_title = entries[i] # 累计总数据用于计算平均值 total_score += current_score total_comms += current_comms # 更新最高得分记录 if current_score > max_score: max_score = current_score max_score_title = current_title # 更新最低得分记录 if current_score < min_score: min_score = current_score min_score_title = current_title # 更新最多评论记录 if current_comms > max_comms: max_comms = current_comms max_comms_title = current_title i += 1 # 计算平均值 avg_score = total_score / entry_count avg_comms = total_comms / entry_count # 输出所有要求的统计结果 print("=====Reddit帖子统计结果=====") print(f"所有帖子平均评论数:{avg_comms:.2f}") print(f"所有帖子平均得分:{avg_score:.2f}") print(f"最高得分对应帖子标题:{max_score_title}(得分:{max_score})") print(f"最低得分对应帖子标题:{min_score_title}(得分:{min_score})") print(f"评论数最多的帖子标题:{max_comms_title}(评论数:{max_comms})") return avg_score with open("reddit_vm.csv", "r", encoding='UTF-8', errors="ignore") as input: entries = [(e['id'], int(e['score']), int(e['comms_num']), e['title']) for e in csv.DictReader(input)] avgScore = analyze(entries)
逻辑说明
- 累计类统计变量初始值设为0,极值类变量默认取第一条数据的值,避免初始值偏差
- 用索引变量
i控制while循环遍历所有条目,每次循环同步更新所有统计值 - 平均值保留2位小数输出,可根据需求调整精度
内容的提问来源于stack exchange,提问作者Spencer
相关产品推荐
相关产品推荐

