You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python打开JSON文件时遭遇JSONDecodeError报错求助

解决读取tweets.json时的JSONDecodeError问题

问题根源

报错提示Expecting ':' delimiter,说明你的tweets.json文件不符合标准JSON格式:要么是单行JSON里的键值对缺失冒号、引号不匹配等语法错误;要么是文件里塞了多个独立的JSON对象(比如每条Tweet占一行,但没做数组包裹或正确分隔),直接用json.load读取就会失败。

可行解决方案

1. 定位修复语法错误

根据报错位置(第1行第8388729列),用VS Code、Notepad++这类编辑器打开文件,跳转到对应位置,检查附近的JSON结构:

  • 确认键是否用双引号包裹(JSON要求键必须是双引号格式)
  • 检查键和值之间有没有漏写冒号
  • 排查是否有多余或缺失的逗号、括号

修复后再用原代码测试。

2. 处理多行独立JSON对象

如果你的文件是每行一个Tweet的JSON(常见的Twitter导出格式),改成逐行读取解析:

import json

tweets = []
with open('tweets.json', 'r', encoding='utf-8') as jfile:
    for line in jfile:
        line = line.strip()
        if not line:
            continue
        try:
            tweet = json.loads(line)
            tweets.append(tweet)
        except json.JSONDecodeError as e:
            print(f"跳过错误行: {e}")

3. 修复拼接错误的单行多JSON

如果文件是把多个JSON对象直接拼在一行(没有用[]包裹,也没加逗号分隔),可以用脚本修复:

import json

# 读取原文件内容
with open('tweets.json', 'r', encoding='utf-8') as f:
    raw_content = f.read()

# 假设每个JSON对象以{"id"开头(根据你的数据特征调整)
split_marker = '{"id"'
parts = raw_content.split(split_marker)
# 重新拼接成合法的JSON数组
fixed_content = '[' + f',{split_marker}'.join(parts[1:]) + ']'

# 写入修复后的文件
with open('fixed_tweets.json', 'w', encoding='utf-8') as f:
    f.write(fixed_content)

# 读取修复后的文件
with open('fixed_tweets.json', 'r', encoding='utf-8') as jfile:
    d = json.load(jfile)

注意:如果你的JSON对象不是以{"id"开头,要换成实际的起始特征字符串。

4. 用容错库解析损坏的JSON

如果手动修复麻烦,试试demjson库,它能解析有小语法错误的JSON:

  1. 先安装:pip install demjson
  2. 读取代码:
import demjson

with open('tweets.json', 'r', encoding='utf-8') as jfile:
    content = jfile.read()
    d = demjson.decode(content)

内容的提问来源于stack exchange,提问作者Kishore Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 10:00:56