如何在Python中处理超大型CSV文件,按user_a获取最大score的user_b
处理大CSV文件的低内存解决方案
针对你的需求——在无法使用Pandas/Dask、避免内存不足的前提下,为每个user_a找到对应最大score的user_b,以下是两种高效低内存的实现方案:
方法1:逐行读取+字典实时更新(内存占用取决于唯一user_a数量)
直接通过Python内置文件迭代器逐行读取,用字典仅记录每个user_a的当前最优结果,无需加载全量数据到内存。
def process_large_csv(input_path, output_path): max_records = {} # 键: user_a,值: (max_score, user_b) with open(input_path, 'r', encoding='utf-8') as infile, \ open(output_path, 'w', encoding='utf-8') as outfile: # 逐行读取,文件对象本身是迭代器,仅加载当前行到内存 for line in infile: line = line.strip() if not line: continue parts = line.split(',') if len(parts) != 3: continue # 跳过格式错误行 # 类型转换,避免字符串占用额外内存 try: user_a = int(parts[0]) user_b = int(parts[1]) score = float(parts[2]) except ValueError: continue # 跳过数据无效行 # 实时更新当前user_a的最优结果 if user_a not in max_records or score > max_records[user_a][0]: max_records[user_a] = (score, user_b) # 写入最终结果 for user_a, (score, user_b) in sorted(max_records.items()): outfile.write(f"{user_a} {user_b} {score}\n") # 调用示例 process_large_csv("input.csv", "output.txt")
该方案的内存占用仅与唯一user_a的数量相关,即使百万行数据,若唯一user_a为几十万级别,内存占用也仅为几十MB,完全可控。
方法2:先排序再逐组处理(内存占用极低,无数据量限制)
如果系统支持外部排序工具(如Linux/macOS的sort命令),先将CSV按user_a排序,使同一user_a的行连续排列,处理时仅需跟踪当前user_a的最优结果,内存占用可忽略不计。
步骤1:系统命令排序(Linux/macOS)
sort -t ',' -k 1n input.csv > sorted_input.csv
参数说明:-t ','指定分隔符为逗号,-k 1n按第一列(user_a)数值排序。
步骤2:Python处理排序后的文件
def process_sorted_csv(input_path, output_path): current_user = None current_max_score = -float('inf') current_best_b = None with open(input_path, 'r', encoding='utf-8') as infile, \ open(output_path, 'w', encoding='utf-8') as outfile: for line in infile: line = line.strip() if not line: continue parts = line.split(',') if len(parts) != 3: continue try: user_a = int(parts[0]) user_b = int(parts[1]) score = float(parts[2]) except ValueError: continue # 初始化第一个用户 if current_user is None: current_user = user_a current_max_score = score current_best_b = user_b # 同一用户,更新最优结果 elif user_a == current_user: if score > current_max_score: current_max_score = score current_best_b = user_b # 切换用户,输出上一个用户的结果并重置 else: outfile.write(f"{current_user} {current_best_b} {current_max_score}\n") current_user = user_a current_max_score = score current_best_b = user_b # 处理最后一个用户 if current_user is not None: outfile.write(f"{current_user} {current_best_b} {current_max_score}\n") # 调用示例 process_sorted_csv("sorted_input.csv", "output.txt")
该方案内存仅存储当前用户的几个变量,无论文件多大都不会出现内存不足问题。
关键内存优化细节
- 绝对避免一次性读取全量文件:不要使用
readlines(),必须用for line in infile的迭代方式,仅加载当前行。 - 最小化数据存储:只保留每个
user_a的最优结果,不存储无关行数据。 - 优化数据类型:将
user_a/user_b转为整数、score转为浮点数,比字符串节省内存。 - 避免结果全量驻留内存:若使用yield生成结果,需逐行写入文件,不要用
list(generator)将所有结果加载到内存。
内容的提问来源于stack exchange,提问作者illuminato
相关产品推荐
相关产品推荐

