Thonny可正常运行的Python代码在Trinket报JSONDecodeError的原因
问题描述
同一段商品比价Python代码在本地离线IDE Thonny中可正常运行,迁移至在线编程平台Trinket运行时,解析JSON环节抛出异常。
核心业务代码逻辑:通过requests拉取两个商品站点的页面内容,用BeautifulSoup解析HTML后,提取type="application/ld+json"的script标签内的结构化数据,调用json.loads()做JSON解析。
对应代码片段:
def compare_prices(product_laughs,product_glomark): html_lau=requests.get(product_laughs).content html_glo=requests.get(product_glomark).content #get the content of the site soup_lau=BeautifulSoup(html_lau,'html.parser') soup_glo=BeautifulSoup(html_glo,'html.parser') glo_products=soup_glo.find("script", type="application/ld+json") glomark_content=glo_products.text #glomark product list print(glomark_content) productList=json.loads(glomark_content)
触发的完整报错信息:
Traceback (most recent call last): File "/tmp/sessions/96aa9a060805dc3c/main.py", line 4, in <module> compare_prices(laughs_coconut,glomark_coconut) File "/tmp/sessions/96aa9a0605dc3c/compare_prices.py", line 27, in compare_prices productList=json.loads(glomark_content) File "/usr/lib/python3.9/json/__init__.py", line 346, in loads return _default_decoder.decode(s) File "/usr/lib/python3.9/json/decoder.py", line 337, in decode obj, end = self.raw_decode(s, idx=_w(s, 0).end()) File "/usr/lib/python3.9/json/decoder.py", line 355, in raw_decode raise JSONDecodeError("Expecting value", s, err.value) from None json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
故障根因
- 报错
Expecting value: line 1 column 1 (char 0)的直接原因是传入json.loads()的参数是空字符串,解析器从第一个字符开始就没找到任何有效JSON内容,和JSON格式错误无关。 - 核心差异来自两个运行环境的网络身份不同:本地Thonny发起请求时用的是个人家庭宽带的出口IP,目标电商站点将其识别为普通用户访问,正常返回完整商品页HTML,代码可以正常定位到ld+json脚本标签、拿到有效结构化数据;Trinket作为公共在线编程平台,出口IP属于云服务商公开IP段,早就被目标站点的反爬系统标记为爬虫流量,返回的要么是空白页、要么是人机验证页、要么是缺失结构化数据的异常页面,导致代码要么找不到目标script标签,要么取到的标签文本为空。
- 代码本身没有做容错处理:默认
find()方法一定能找到目标标签、标签内一定有有效内容,一旦请求被拦截就会直接触发解析报错。
修复方案
- 首先给requests请求添加真实浏览器的User-Agent请求头,模拟普通用户访问,降低被反爬识别的概率。
- 增加空值校验逻辑,不要默认页面元素一定存在,出现拦截时直接抛出明确提示,避免无意义的JSON解析报错。
- 调整后的可运行代码示例:
import requests from bs4 import BeautifulSoup import json # 模拟桌面端Chrome浏览器的请求头 REQUEST_HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8" } def compare_prices(product_laughs,product_glomark): # 发起请求时携带请求头 resp_lau = requests.get(product_laughs, headers=REQUEST_HEADERS, timeout=10) resp_lau.raise_for_status() # 遇到HTTP错误状态码直接抛出异常 html_lau = resp_lau.content resp_glo = requests.get(product_glomark, headers=REQUEST_HEADERS, timeout=10) resp_glo.raise_for_status() html_glo = resp_glo.content soup_lau=BeautifulSoup(html_lau,'html.parser') soup_glo=BeautifulSoup(html_glo,'html.parser') glo_products=soup_glo.find("script", type="application/ld+json") # 校验是否找到目标标签 if not glo_products: raise RuntimeError("当前网络环境被目标站点反爬拦截,页面中未找到商品结构化数据标签") glomark_content = glo_products.text.strip() # 去除首尾空白、换行字符 # 校验标签内容是否为空 if not glomark_content: raise RuntimeError("获取到的结构化数据标签内容为空,请求可能被反爬拦截") productList=json.loads(glomark_content)
- 如果添加请求头后在Trinket平台依然报错,说明Trinket的出口IP已经被目标站点完全封禁,公共在线平台的IP池被大量用户共享用来爬取数据,基本没有稳定绕过反爬的可能,这类场景直接切换到本地环境运行即可。
内容的提问来源于stack exchange,提问作者A_Ruwantha
相关产品推荐
相关产品推荐

