使用Python Scrapy加载JSON时出现JSONDecodeError Extra data错误如何解决
问题排查与解决方案
报错原因
JSONDecodeError: Extra data 报错的核心原因是你切割得到的el字符串不是纯JSON内容,尾部包含了JS脚本的其他多余字符(比如逗号、右括号、分号、后续变量定义等),不符合JSON格式规范,因此无法直接解析。
修复方案
通用解决方案(适配嵌套JSON结构)
通过匹配JSON根节点的闭合括号自动提取纯JSON内容,避免硬编码切割适配性差的问题:
- 实现一个括号匹配的工具函数,自动定位JSON的起止位置
- 提取到纯JSON字符串后再执行解析
修改后的完整代码
import scrapy import json class CodeSpider(scrapy.Spider): name = 'trip' allowed_domains = ['tripadvisor.com'] start_urls = ['https://www.tripadvisor.com/Restaurant_Review-g155019-d5058155-Reviews-The_Tilted_Dog_Pub_Kitchen-Toronto_Ontario.html'] # 括号匹配工具函数,提取纯JSON内容 def extract_json(self, raw_str, root_char='{'): start = raw_str.find(root_char) if start == -1: return '' count = 1 end = start + 1 close_char = '}' if root_char == '{' else ']' while end < len(raw_str) and count > 0: if raw_str[end] == root_char: count += 1 elif raw_str[end] == close_char: count -= 1 end += 1 return raw_str[start:end-1] def start_requests(self): for a in self.start_urls: yield scrapy.Request(url = a,callback= self.parse_data) def parse_data(self, response): script_text = response.xpath('//div[@id="taplc_footer_js_globals_0"]/following::script[1]/text()').get() raw_el = script_text.split('"responses":')[1] # 提取纯JSON,如果responses对应的值是数组,把root_char参数改成'['即可 pure_json = self.extract_json(raw_el, root_char='{') data = json.loads(pure_json) # 后续业务处理逻辑
快速调试方案
如果需要临时验证效果,可先将切割后的raw_el写入本地文件查看冗余内容,手动裁剪尾部多余字符:
# 写入本地文件查看完整内容,定位冗余部分 with open('raw_content.txt', 'w', encoding='utf-8') as f: f.write(raw_el) # 示例:如果尾部多余内容是`, "otherConfig": xxx});`,可手动裁剪 pure_json = raw_el.split(', "otherConfig"')[0]
内容的提问来源于stack exchange,提问作者Haris Ahmad
相关产品推荐
相关产品推荐

