使用pd.read_json读取JSONL文件时触发AttributeError错误求助
问题解决:AttributeError: 'list' object has no attribute 'items'
错误原因
你在使用pd.read_json()时同时指定了lines=True和orient='split',这两个参数的适用场景完全冲突:
lines=True用于解析JSONL格式文件(每行一个独立的JSON对象),解析后得到的是JSON对象组成的列表。orient='split'要求JSON结构是包含columns、data、index键的字典,pandas会尝试对解析结果调用items()方法,但列表没有这个属性,因此触发错误。
解决方案
直接移除orient='split'参数,lines=True已经足够适配你的JSONL文件:
import pandas as pd filename = "/content/CapybaraPure_Decontaminated.jsonl" instruction_dataset_df = pd.read_json(filename, lines=True) examples = instruction_dataset_df.to_dict()
补充说明
如果你的文件确实是split格式(而非JSONL),那么文件结构应该类似这样:
{ "columns": ["instruction", "response"], "data": [["问:什么是熊猫?", "答:熊猫是中国国宝..."], ["问:天空为什么是蓝的?", "答:因为瑞利散射..."]], "index": [0, 1] }
这种情况下才需要使用orient='split',同时不能加lines=True。
内容的提问来源于stack exchange,提问作者davinci
相关产品推荐
相关产品推荐

