Python如何从JSON文件提取指定text字段内容并合并生成txt文件
JSON Lines格式text字段提取合并方案
处理逻辑确认
是的,你当前使用的每行存储一个独立JSON对象的格式为JSON Lines格式,无法直接将整个文件作为单个合法JSON加载,必须逐行遍历文件内容完成提取操作。
常规实现代码(适合中小文件)
直接将所有提取的text存入内存后合并写入,逻辑简单清晰,适配你给出的需求效果:
import json # 替换为你的实际文件路径 input_file = "source.jsonl" output_file = "merged_text.txt" text_collection = [] # 逐行读取源文件 with open(input_file, "r", encoding="utf-8") as f: for line in f: stripped_line = line.strip() # 跳过空行 if not stripped_line: continue # 解析单条JSON并提取text json_data = json.loads(stripped_line) text_collection.append(json_data["text"]) # 用空格拼接所有内容(与你提供的示例输出效果一致) merged_content = " ".join(text_collection) # 写入输出文件 with open(output_file, "w", encoding="utf-8") as f: f.write(merged_content)
大文件优化实现(适合GB级以上文件)
如果源文件体积很大,为了避免占用过多内存,可以逐行提取后直接写入输出文件,无需暂存所有text内容:
import json input_file = "source.jsonl" output_file = "merged_text.txt" with open(input_file, "r", encoding="utf-8") as f_in, open(output_file, "w", encoding="utf-8") as f_out: # 标记是否为第一条内容,避免开头多输出空格 is_first = True for line in f_in: stripped_line = line.strip() if not stripped_line: continue try: json_data = json.loads(stripped_line) if "text" in json_data: if not is_first: f_out.write(" ") f_out.write(json_data["text"]) is_first = False except json.JSONDecodeError: # 可自行添加错误日志打印逻辑 continue
容错说明
上述优化版代码已经添加了基础异常捕获,可跳过格式错误的JSON行、以及缺失text字段的行,避免程序意外中断。
内容的提问来源于stack exchange,提问作者futuredataengineer
相关产品推荐
相关产品推荐

