You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中处理超大型CSV文件,按user_a获取最大score的user_b

处理大CSV文件的低内存解决方案

针对你的需求——在无法使用Pandas/Dask、避免内存不足的前提下,为每个user_a找到对应最大score的user_b,以下是两种高效低内存的实现方案:

方法1:逐行读取+字典实时更新(内存占用取决于唯一user_a数量)

直接通过Python内置文件迭代器逐行读取,用字典仅记录每个user_a的当前最优结果,无需加载全量数据到内存。

def process_large_csv(input_path, output_path):
    max_records = {}  # 键: user_a,值: (max_score, user_b)
    with open(input_path, 'r', encoding='utf-8') as infile, \
         open(output_path, 'w', encoding='utf-8') as outfile:
        # 逐行读取,文件对象本身是迭代器,仅加载当前行到内存
        for line in infile:
            line = line.strip()
            if not line:
                continue
            parts = line.split(',')
            if len(parts) != 3:
                continue  # 跳过格式错误行
            # 类型转换,避免字符串占用额外内存
            try:
                user_a = int(parts[0])
                user_b = int(parts[1])
                score = float(parts[2])
            except ValueError:
                continue  # 跳过数据无效行
            # 实时更新当前user_a的最优结果
            if user_a not in max_records or score > max_records[user_a][0]:
                max_records[user_a] = (score, user_b)
        # 写入最终结果
        for user_a, (score, user_b) in sorted(max_records.items()):
            outfile.write(f"{user_a} {user_b} {score}\n")

# 调用示例
process_large_csv("input.csv", "output.txt")

该方案的内存占用仅与唯一user_a的数量相关,即使百万行数据,若唯一user_a为几十万级别,内存占用也仅为几十MB,完全可控。

方法2:先排序再逐组处理(内存占用极低,无数据量限制)

如果系统支持外部排序工具(如Linux/macOS的sort命令),先将CSV按user_a排序,使同一user_a的行连续排列,处理时仅需跟踪当前user_a的最优结果,内存占用可忽略不计。

步骤1:系统命令排序(Linux/macOS)

sort -t ',' -k 1n input.csv > sorted_input.csv

参数说明:-t ','指定分隔符为逗号,-k 1n按第一列(user_a)数值排序。

步骤2:Python处理排序后的文件

def process_sorted_csv(input_path, output_path):
    current_user = None
    current_max_score = -float('inf')
    current_best_b = None
    with open(input_path, 'r', encoding='utf-8') as infile, \
         open(output_path, 'w', encoding='utf-8') as outfile:
        for line in infile:
            line = line.strip()
            if not line:
                continue
            parts = line.split(',')
            if len(parts) != 3:
                continue
            try:
                user_a = int(parts[0])
                user_b = int(parts[1])
                score = float(parts[2])
            except ValueError:
                continue
            # 初始化第一个用户
            if current_user is None:
                current_user = user_a
                current_max_score = score
                current_best_b = user_b
            # 同一用户,更新最优结果
            elif user_a == current_user:
                if score > current_max_score:
                    current_max_score = score
                    current_best_b = user_b
            # 切换用户,输出上一个用户的结果并重置
            else:
                outfile.write(f"{current_user} {current_best_b} {current_max_score}\n")
                current_user = user_a
                current_max_score = score
                current_best_b = user_b
        # 处理最后一个用户
        if current_user is not None:
            outfile.write(f"{current_user} {current_best_b} {current_max_score}\n")

# 调用示例
process_sorted_csv("sorted_input.csv", "output.txt")

该方案内存仅存储当前用户的几个变量,无论文件多大都不会出现内存不足问题。

关键内存优化细节

  1. 绝对避免一次性读取全量文件:不要使用readlines(),必须用for line in infile的迭代方式,仅加载当前行。
  2. 最小化数据存储:只保留每个user_a的最优结果,不存储无关行数据。
  3. 优化数据类型:将user_a/user_b转为整数、score转为浮点数,比字符串节省内存。
  4. 避免结果全量驻留内存:若使用yield生成结果,需逐行写入文件,不要用list(generator)将所有结果加载到内存。

内容的提问来源于stack exchange,提问作者illuminato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 13:10:23