You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量循环读写并修改多个序列命名的JSON文件?

批量处理文件的实现方案

1. 基于Python的批量处理代码

假设你已经有处理单个文件的逻辑,只需在外层添加循环遍历所有目标文件即可,直接复用你的单文件处理逻辑:

import json

def process_single_file(input_file, output_file):
    # 这里放入你已实现的单个文件处理代码
    with open(input_file, 'r', encoding='utf-8') as f_in:
        with open(output_file, 'w', encoding='utf-8') as f_out:
            for line in f_in:
                try:
                    data = json.loads(line.strip())
                    # 仅替换首个"play "为空字符串
                    if 'utterance' in data and isinstance(data['utterance'], str):
                        data['utterance'] = data['utterance'].replace('play ', '', 1)
                    # 写入处理后的JSON行
                    f_out.write(json.dumps(data, ensure_ascii=False) + '\n')
                except json.JSONDecodeError:
                    # 遇到无效JSON行时直接原样写入(可根据需求调整逻辑)
                    f_out.write(line)

# 批量遍历所有16个文件
for file_index in range(16):
    # 格式化数字为5位带前导零的字符串,匹配文件名规则
    num_str = f"{file_index:05d}"
    input_path = f"test-{num_str}-of-00016"
    output_path = f"file-{num_str}-of-00016.txt"
    
    print(f"处理中: {input_path} → {output_path}")
    process_single_file(input_path, output_path)

print("所有文件处理完成!")

2. 关键细节说明

  • 文件名匹配:用range(16)生成0到15的索引,通过f"{file_index:05d}"自动补前导零,精准匹配你需要的文件名格式。
  • 逻辑复用:将你已有的单文件处理代码直接放入process_single_file函数,无需修改核心处理逻辑。
  • 容错处理:添加了JSON解码异常捕获,避免单个无效行导致整个文件处理中断,可根据实际需求调整(比如记录错误日志)。

3. 运行方式

把代码保存为batch_process.py,放在目标文件所在目录,直接执行:

python batch_process.py

内容的提问来源于stack exchange,提问作者vkaul11

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 01:15:53