如何批量循环读写并修改多个序列命名的JSON文件?
批量处理文件的实现方案
1. 基于Python的批量处理代码
假设你已经有处理单个文件的逻辑,只需在外层添加循环遍历所有目标文件即可,直接复用你的单文件处理逻辑:
import json def process_single_file(input_file, output_file): # 这里放入你已实现的单个文件处理代码 with open(input_file, 'r', encoding='utf-8') as f_in: with open(output_file, 'w', encoding='utf-8') as f_out: for line in f_in: try: data = json.loads(line.strip()) # 仅替换首个"play "为空字符串 if 'utterance' in data and isinstance(data['utterance'], str): data['utterance'] = data['utterance'].replace('play ', '', 1) # 写入处理后的JSON行 f_out.write(json.dumps(data, ensure_ascii=False) + '\n') except json.JSONDecodeError: # 遇到无效JSON行时直接原样写入(可根据需求调整逻辑) f_out.write(line) # 批量遍历所有16个文件 for file_index in range(16): # 格式化数字为5位带前导零的字符串,匹配文件名规则 num_str = f"{file_index:05d}" input_path = f"test-{num_str}-of-00016" output_path = f"file-{num_str}-of-00016.txt" print(f"处理中: {input_path} → {output_path}") process_single_file(input_path, output_path) print("所有文件处理完成!")
2. 关键细节说明
- 文件名匹配:用
range(16)生成0到15的索引,通过f"{file_index:05d}"自动补前导零,精准匹配你需要的文件名格式。 - 逻辑复用:将你已有的单文件处理代码直接放入
process_single_file函数,无需修改核心处理逻辑。 - 容错处理:添加了JSON解码异常捕获,避免单个无效行导致整个文件处理中断,可根据实际需求调整(比如记录错误日志)。
3. 运行方式
把代码保存为batch_process.py,放在目标文件所在目录,直接执行:
python batch_process.py
内容的提问来源于stack exchange,提问作者vkaul11
相关产品推荐
相关产品推荐

