You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python通过正则表达式从CSV文件提取PIN值至新CSV

从大容量CSV中提取17位PIN值并写入新CSV

你的代码存在几个问题:一次性读入大文件会导致内存溢出、正则匹配绕了弯路、还有拼写错误(re.finall应该是re.findall),而且缺少写入文件的逻辑。下面是优化后的实现,专门针对大容量文件设计:

import re

# 输入文件路径
input_path = r"/Users/sj4/Downloads/python_work/Rabbiterror.csv"
# 输出文件路径
output_path = r"/Users/sj4/Downloads/python_work/extracted_pins.csv"

# 精准匹配"pin":"后面的17位字符,用捕获组直接提取PIN
pin_pattern = re.compile(r'"pin":"(\w{17})"')

# 逐行读取大文件,避免内存占用过高
with open(input_path, 'r', encoding='utf-8') as infile, open(output_path, 'w', encoding='utf-8') as outfile:
    for line in infile:
        # 查找当前行所有符合的PIN
        pins = pin_pattern.findall(line)
        for pin in pins:
            # 每个PIN占一行写入输出文件
            outfile.write(f"{pin}\n")

关键说明:

  • 逐行处理:不再一次性把整个文件读入内存,适合GB级别的大容量CSV
  • 精准正则:r'"pin":"(\w{17})"'直接定位到pin字段,通过捕获组直接拿到17位PIN值,无需二次匹配
  • 文件安全:使用with语句自动管理文件句柄,避免忘记关闭文件导致的资源泄漏
  • 可选去重:如果需要去除重复的PIN,可以先把PIN存入集合再写入,比如:
    seen_pins = set()
    with open(input_path, 'r', encoding='utf-8') as infile, open(output_path, 'w', encoding='utf-8') as outfile:
        for line in infile:
            pins = pin_pattern.findall(line)
            for pin in pins:
                if pin not in seen_pins:
                    seen_pins.add(pin)
                    outfile.write(f"{pin}\n")
    

内容的提问来源于stack exchange,提问作者Shenk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 14:46:01