You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3.6加载含多条推文的JSON文件遇ValueError报错求助

解决JSON加载时的ValueError: Extra Data问题

嘿,这个错误我太熟悉了!问题根源在于你的JSON文件结构不符合Python json模块的解析要求。

错误原因

你提供的JSON片段是独立的单条推文对象,而整个文件里应该是多条这样的对象直接拼接在一起(没有用数组包裹,也没在对象间加逗号分隔)。Python的json.load()默认只能解析单个JSON结构——要么是一个单独的{}对象,要么是一个包含所有对象的[]数组。当它解析完第一个推文对象后,发现后面还有额外的JSON数据,就会抛出Extra data的错误。

解决方法

根据你的情况,有两种常用的处理方式:

方式1:修改源JSON文件(推荐,如果你能编辑文件的话)

把所有推文对象包裹在一个数组里,每个对象之间用逗号分隔,格式如下:

[
  {"contributors": null, "coordinates": null, ...},
  {"contributors": null, "coordinates": null, ...},
  ...
]

修改完成后,直接用常规的加载代码就能正常读取成一个推文列表:

import json

with open('tweets.json', 'r', encoding='utf-8') as f:
    tweets = json.load(f)

方式2:在Python中逐段解析(无法修改源文件时)

如果你的JSON文件是每行一条推文(这是Twitter导出数据的常见格式),可以逐行读取并解析:

import json

tweets = []
with open('tweets.json', 'r', encoding='utf-8') as f:
    for line in f:
        # 跳过空行避免解析错误
        cleaned_line = line.strip()
        if cleaned_line:
            tweet = json.loads(cleaned_line)
            tweets.append(tweet)

如果你的文件里的JSON是连在一起没有换行的,就用JSONDecoder的raw_decode()方法来逐个提取对象:

import json

tweets = []
decoder = json.JSONDecoder()
with open('tweets.json', 'r', encoding='utf-8') as f:
    raw_data = f.read()
    current_pos = 0
    total_length = len(raw_data)
    
    while current_pos < total_length:
        try:
            # 从当前位置解析出一个JSON对象,返回对象和下一个起始位置
            tweet, current_pos = decoder.raw_decode(raw_data, current_pos)
            tweets.append(tweet)
        except json.JSONDecodeError:
            # 遇到无效字符(比如空格、换行)时,跳过一个字符继续
            current_pos += 1

内容的提问来源于stack exchange,提问作者Charbel Hanna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:09:50