如何在Python中不使用pandas找出CSV列中最频繁的字符串值
不依赖pandas找出CSV列中最频繁且首次出现的值
下面是针对需求的完整实现,完全用Python内置模块完成,适合新手理解:
import csv def find_most_frequent_height(csv_file_path): # 初始化两个字典:一个存计数,一个存每个值首次出现的行索引 height_counts = {} first_occurrence = {} current_row = 0 with open(csv_file_path, 'r', newline='', encoding='utf-8') as file: # 用DictReader按列名读取,直接通过列名获取数据 reader = csv.DictReader(file) for row in reader: height = row['height_c'] # 只处理0-10的字符串(符合题目要求) if height not in [str(i) for i in range(11)]: continue # 跳过不符合格式的值,可选操作 # 如果是第一次出现,记录行号 if height not in first_occurrence: first_occurrence[height] = current_row # 更新计数 height_counts[height] = height_counts.get(height, 0) + 1 current_row += 1 # 处理空文件或无有效数据的情况 if not height_counts: return '' # 或者根据需求返回默认值 # 按出现次数降序,次数相同则按首次出现的行号升序排序 sorted_heights = sorted(height_counts.items(), key=lambda x: (-x[1], first_occurrence[x[0]])) # 返回第一个元素的键(即最频繁且首次出现的值) return sorted_heights[0][0] # 调用示例 if __name__ == '__main__': result = find_most_frequent_height('personal_info.csv') print(f"最频繁的height_c值是: {result}")
代码说明:
- 用内置
csv.DictReader读取文件,不需要手动处理列索引,直接通过row['height_c']获取目标列数据,对新手友好 - 用
height_counts统计每个值的出现次数,first_occurrence记录每个值第一次出现的行号,解决"次数相同返回首次出现"的需求 - 排序时用
lambda表达式:先按次数降序(-x[1]),次数相同则按首次出现的行号升序(first_occurrence[x[0]]),确保取到最早出现的高频值 - 加入了对非0-10字符串的过滤逻辑,可选保留,避免无效数据干扰统计
测试验证
假设你的personal_info.csv内容如下:
full_name,weight_b,height_c Alice,55,2 Bob,60,3 Charlie,70,2 David,65,4 Eve,58,2 Frank,72,3
运行代码后会返回'2',符合示例预期。
内容的提问来源于stack exchange,提问作者Emppy
相关产品推荐
相关产品推荐

