使用Python打开JSON文件时遭遇JSONDecodeError报错求助
解决读取tweets.json时的JSONDecodeError问题
问题根源
报错提示Expecting ':' delimiter,说明你的tweets.json文件不符合标准JSON格式:要么是单行JSON里的键值对缺失冒号、引号不匹配等语法错误;要么是文件里塞了多个独立的JSON对象(比如每条Tweet占一行,但没做数组包裹或正确分隔),直接用json.load读取就会失败。
可行解决方案
1. 定位修复语法错误
根据报错位置(第1行第8388729列),用VS Code、Notepad++这类编辑器打开文件,跳转到对应位置,检查附近的JSON结构:
- 确认键是否用双引号包裹(JSON要求键必须是双引号格式)
- 检查键和值之间有没有漏写冒号
- 排查是否有多余或缺失的逗号、括号
修复后再用原代码测试。
2. 处理多行独立JSON对象
如果你的文件是每行一个Tweet的JSON(常见的Twitter导出格式),改成逐行读取解析:
import json tweets = [] with open('tweets.json', 'r', encoding='utf-8') as jfile: for line in jfile: line = line.strip() if not line: continue try: tweet = json.loads(line) tweets.append(tweet) except json.JSONDecodeError as e: print(f"跳过错误行: {e}")
3. 修复拼接错误的单行多JSON
如果文件是把多个JSON对象直接拼在一行(没有用[]包裹,也没加逗号分隔),可以用脚本修复:
import json # 读取原文件内容 with open('tweets.json', 'r', encoding='utf-8') as f: raw_content = f.read() # 假设每个JSON对象以{"id"开头(根据你的数据特征调整) split_marker = '{"id"' parts = raw_content.split(split_marker) # 重新拼接成合法的JSON数组 fixed_content = '[' + f',{split_marker}'.join(parts[1:]) + ']' # 写入修复后的文件 with open('fixed_tweets.json', 'w', encoding='utf-8') as f: f.write(fixed_content) # 读取修复后的文件 with open('fixed_tweets.json', 'r', encoding='utf-8') as jfile: d = json.load(jfile)
注意:如果你的JSON对象不是以{"id"开头,要换成实际的起始特征字符串。
4. 用容错库解析损坏的JSON
如果手动修复麻烦,试试demjson库,它能解析有小语法错误的JSON:
- 先安装:
pip install demjson - 读取代码:
import demjson with open('tweets.json', 'r', encoding='utf-8') as jfile: content = jfile.read() d = demjson.decode(content)
内容的提问来源于stack exchange,提问作者Kishore Kumar
相关产品推荐
相关产品推荐

