如何将Pandas DataFrame转换为指定结构的嵌套JSONL文件
问题根源
你遇到的问题核心有两点:
quantity details列当前是字符串类型,不是Python字典结构,导出JSON时会被整体识别为字符串,不会自动转成嵌套对象- 手动拼接JSON字符串、替换引号的操作会触发嵌套引号的转义错误,导致输出的JSON格式异常
解决方案
方案1:基于你现有的quantity details字符串列处理
import pandas as pd import json # 第一步:将quantity details列从字符串转为字典 df["quantity details"] = df["quantity details"].apply( lambda x: json.loads(f"{{{x}}}") ) # 第二步:逐行构造符合要求的记录,写入JSONL文件 with open("file_name_here.jsonl", "w", encoding="utf-8") as f: for _, row in df.iterrows(): record = {"input": row.to_dict()} # 用json库内置方法序列化,自动处理转义,避免格式错误 f.write(json.dumps(record, ensure_ascii=False) + "\n")
方案2(更推荐):生成quantity details列时直接存字典
如果你是从原始的quantity、locationId两列合并得到quantity details,建议直接生成字典格式的列,省去字符串转字典的步骤,出错概率更低:
import pandas as pd import json # 直接生成字典类型的quantity details列 df["quantity details"] = df.apply( lambda row: { "availableQuantity": row["quantity"], "locationId": row["locationId"] }, axis=1 ) # 后续写入JSONL的逻辑和方案1一致 with open("file_name_here.jsonl", "w", encoding="utf-8") as f: for _, row in df.iterrows(): record = {"input": row.to_dict()} f.write(json.dumps(record, ensure_ascii=False) + "\n")
内容的提问来源于stack exchange,提问作者JJasperr
相关产品推荐
相关产品推荐

