You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何从JSON文件提取指定text字段内容并合并生成txt文件

JSON Lines格式text字段提取合并方案

处理逻辑确认

是的,你当前使用的每行存储一个独立JSON对象的格式为JSON Lines格式,无法直接将整个文件作为单个合法JSON加载,必须逐行遍历文件内容完成提取操作。

常规实现代码(适合中小文件)

直接将所有提取的text存入内存后合并写入,逻辑简单清晰,适配你给出的需求效果:

import json

# 替换为你的实际文件路径
input_file = "source.jsonl"
output_file = "merged_text.txt"

text_collection = []

# 逐行读取源文件
with open(input_file, "r", encoding="utf-8") as f:
    for line in f:
        stripped_line = line.strip()
        # 跳过空行
        if not stripped_line:
            continue
        # 解析单条JSON并提取text
        json_data = json.loads(stripped_line)
        text_collection.append(json_data["text"])

# 用空格拼接所有内容(与你提供的示例输出效果一致)
merged_content = " ".join(text_collection)

# 写入输出文件
with open(output_file, "w", encoding="utf-8") as f:
    f.write(merged_content)

大文件优化实现(适合GB级以上文件)

如果源文件体积很大,为了避免占用过多内存,可以逐行提取后直接写入输出文件,无需暂存所有text内容:

import json

input_file = "source.jsonl"
output_file = "merged_text.txt"

with open(input_file, "r", encoding="utf-8") as f_in, open(output_file, "w", encoding="utf-8") as f_out:
    # 标记是否为第一条内容,避免开头多输出空格
    is_first = True
    for line in f_in:
        stripped_line = line.strip()
        if not stripped_line:
            continue
        try:
            json_data = json.loads(stripped_line)
            if "text" in json_data:
                if not is_first:
                    f_out.write(" ")
                f_out.write(json_data["text"])
                is_first = False
        except json.JSONDecodeError:
            # 可自行添加错误日志打印逻辑
            continue

容错说明

上述优化版代码已经添加了基础异常捕获,可跳过格式错误的JSON行、以及缺失text字段的行,避免程序意外中断。

内容的提问来源于stack exchange,提问作者futuredataengineer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 22:15:03