You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

对比两个大型CSV文件(8GB&5GB),哪种方法更高效提取不在file1的ID?

大型CSV文件ID对比方案

需求:对比两个分别为8GB和5GB的单列GUID格式CSV文件,输出所有不存在于file1中的ID。

方法1:使用Pandas实现

借助Pandas的合并与筛选功能完成对比,代码逻辑清晰,但需注意大文件加载可能带来较高内存占用:

df = pd.read_csv(file)
df1 = pd.read_csv(file1)

df = df.merge(df1, on=['id'], how="outer", indicator=True).query('_merge=="left_only"')
df['id'].to_csv(output_path, index=False)

方法2:使用原生Python+集合实现

利用集合O(1)查询的特性优化性能,通过逐行读取文件降低内存消耗,更适合处理超大型文件:

with open(file1, 'r') as t1:
    file = set(t1)

with open(file, 'r') as t2, open(output_path, 'w') as outFile:
    for line in t2:
        if line not in file:
            outFile.write(line)

内容的提问来源于stack exchange,提问作者jonas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 09:35:22