You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将分段结构化文本转为CSV?Notepad++转置或Python更合适?

用Python还是Notepad++处理批量文本转CSV?

肯定是Python更合适,原因如下:

  • 数据规模适配:13.2万行的文本量不算小,Notepad++的转置功能依赖界面编辑,处理大文件很容易卡顿甚至直接崩溃;Python可以流式读取文件,逐行处理,内存占用低,能稳定搞定这种量级的数据。
  • 可变字段的灵活处理:你的数据里大部分记录字段一致,但部分有额外字段(比如第三条的Limits)。Notepad++的转置是静态的行/列转换,没法自动识别字段差异,更没法给缺失字段统一补n/a;Python可以用字典动态收集每条记录的字段,最后统一对齐你需要的列,缺失值自动填充默认值,完美适配这种不规则的记录格式。
  • 效率与复用性:用Python写个脚本,一次写完以后类似需求直接改改参数就能用;Notepad++每次都得手动操作,不仅麻烦,还容易出错,重复劳动效率极低。

给你个简单的Python实现示例,直接就能用:

import csv

# 定义最终CSV要输出的列
output_columns = ['name', 'place', 'initial', 'limit']

with open('你的输入文件名.txt', 'r', encoding='utf-8') as input_file, open('输出结果.csv', 'w', newline='', encoding='utf-8') as output_file:
    # 初始化CSV写入器
    csv_writer = csv.DictWriter(output_file, fieldnames=output_columns)
    csv_writer.writeheader()
    
    current_record = {}
    for line in input_file:
        cleaned_line = line.strip()
        # 遇到分隔符时,处理当前记录并写入
        if cleaned_line == '********':
            if current_record:
                # 给缺失的字段补n/a
                for col in output_columns:
                    current_record.setdefault(col, 'n/a')
                # 按指定列顺序整理数据并写入
                csv_writer.writerow({col: current_record[col] for col in output_columns})
                current_record = {}
        else:
            # 解析每行的字段和值
            if ' : ' in cleaned_line:
                key, value = cleaned_line.split(' : ', 1)
                # 字段名转小写,统一格式
                current_record[key.lower()] = value
    # 处理文件末尾最后一条记录
    if current_record:
        for col in output_columns:
            current_record.setdefault(col, 'n/a')
        csv_writer.writerow({col: current_record[col] for col in output_columns})

内容的提问来源于stack exchange,提问作者Gene Parmesan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 03:52:27