You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何生成符合GPT-3微调要求的JSON Lines格式数据集?

生成GPT-3微调所需的JSON Lines格式数据集

要生成符合OpenAI微调要求的JSON Lines格式文件,你需要调整代码逻辑,确保每行写入一个独立的prompt-completion JSON对象,而非将整个列表一次性序列化。

修改后的代码

import json

# 存储所有训练数据的列表
training_data = [
    {"prompt": "<text1>", "completion": "<text to be generated1>"},
    {"prompt": "<text2>", "completion": "<text to be generated2>"}
]

# 写入JSON Lines格式文件(建议用.jsonl后缀明确格式)
with open("sample2.jsonl", "w") as outfile:
    for item in training_data:
        # 逐个序列化单个JSON对象
        json.dump(item, outfile)
        # 每个对象后添加换行符,满足JSON Lines规范
        outfile.write("\n")

代码说明

  • 先将所有prompt-completion对存入列表,便于遍历处理;
  • 遍历列表时,用json.dump()单独写入每个对象;
  • 手动添加换行符,保证每行是一个独立的JSON对象,这是JSON Lines格式的核心要求。

验证读取结果

可以用以下代码确认生成的文件格式是否正确:

import json

parsed_data = []
with open('sample2.jsonl', 'r') as openfile:
    # 逐行读取并解析每个JSON对象
    for line in openfile:
        parsed_item = json.loads(line)
        parsed_data.append(parsed_item)

print(parsed_data)
print(type(parsed_data))

执行后会得到包含两个字典的列表,同时文件本身的内容是每行一个独立的JSON对象,完全符合OpenAI微调的数据集格式要求。

内容的提问来源于stack exchange,提问作者John Angelopoulos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 17:01:08