You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从3亿条记录的大文件中快速提取指定邮箱?

高效提取目标邮箱记录的方法

一、用GNU grep(Linux/macOS/WSL适用)

这是最快的方案之一,grep专门针对文本匹配做了底层优化,处理大文件比编辑器高效得多:

  • 保持目标邮箱文件target_emails.txt为每行一个邮箱的格式(不用转成|分隔)
  • 执行命令:
grep -F -f target_emails.txt large_file.txt > result.txt

参数说明:

  • -F:把每行目标邮箱当作固定字符串匹配,跳过正则解析,速度翻倍
  • -f:从target_emails.txt读取所有匹配规则
  • 结果会直接写入result.txt,全程不用加载整个大文件到内存

如果怕行内有类似邮箱的干扰内容,比如a@gmail.com.cn被误匹配成a@gmail.com,可以加-w参数匹配整词(前提是邮箱周围有空格/标点这类分隔符):

grep -F -w -f target_emails.txt large_file.txt > result.txt

二、用awk脚本(适合结构化文本)

如果你的大文件每行是结构化数据(比如用空格/逗号分隔字段),awk的效率也很高,还能精准定位字段:

场景1:邮箱是固定字段(比如第2个字段,空格分隔)

awk -F ' ' 'NR==FNR{emails[$0]=1; next} $2 in emails' target_emails.txt large_file.txt > result.txt

场景2:邮箱位置不固定,整行包含即可

awk 'NR==FNR{emails[$0]=1; next} {for(e in emails) if(index($0,e)) {print; break}}' target_emails.txt large_file.txt > result.txt

原理:先把目标邮箱存入数组,再逐行扫描大文件,匹配到就输出,内存占用极低。

三、Python脚本(跨平台,Windows也能用)

如果没有Linux环境,用Python写个小脚本也能高效处理,关键是逐行读取,避免加载整个3亿行文件到内存:

def extract_matching_records(target_path, large_path, output_path):
    # 把目标邮箱存进集合,查找速度是O(1)
    target_emails = set()
    with open(target_path, 'r', encoding='utf-8') as f:
        for line in f:
            email = line.strip()
            if email:
                target_emails.add(email)
    
    # 逐行处理大文件,匹配就写入结果
    with open(large_path, 'r', encoding='utf-8') as infile, open(output_path, 'w', encoding='utf-8') as outfile:
        for line in infile:
            # 如果邮箱是固定字段,比如逗号分隔的第3个字段,替换成下面的代码:
            # parts = line.strip().split(',')
            # if len(parts)>=3 and parts[2] in target_emails:
            if any(email in line for email in target_emails):
                outfile.write(line)

if __name__ == '__main__':
    # 替换成你的实际文件路径
    extract_matching_records("target_emails.txt", "large_file.txt", "result.txt")

为啥之前的方法慢?

Emeditor是文本编辑器,核心功能是编辑而非批量处理超大型文件。把邮箱拼成|分隔的正则模式,正则匹配本身就比固定字符串匹配慢,再加上编辑器要把3亿行内容加载到内存,内存压力大导致速度暴跌,完全不是这类场景的最优解。

内容的提问来源于stack exchange,提问作者Marcin Tomczak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 13:42:33