You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Thonny可正常运行的Python代码在Trinket报JSONDecodeError的原因

问题描述

同一段商品比价Python代码在本地离线IDE Thonny中可正常运行,迁移至在线编程平台Trinket运行时,解析JSON环节抛出异常。
核心业务代码逻辑:通过requests拉取两个商品站点的页面内容,用BeautifulSoup解析HTML后,提取type="application/ld+json"的script标签内的结构化数据,调用json.loads()做JSON解析。
对应代码片段:

def compare_prices(product_laughs,product_glomark):
    html_lau=requests.get(product_laughs).content 
    html_glo=requests.get(product_glomark).content #get the content of the site

    soup_lau=BeautifulSoup(html_lau,'html.parser')
    soup_glo=BeautifulSoup(html_glo,'html.parser')
    
    glo_products=soup_glo.find("script", type="application/ld+json")
    glomark_content=glo_products.text #glomark product list
    print(glomark_content)
    productList=json.loads(glomark_content)

触发的完整报错信息:

Traceback (most recent call last):
  File "/tmp/sessions/96aa9a060805dc3c/main.py", line 4, in <module>
    compare_prices(laughs_coconut,glomark_coconut)
  File "/tmp/sessions/96aa9a0605dc3c/compare_prices.py", line 27, in compare_prices
    productList=json.loads(glomark_content)
  File "/usr/lib/python3.9/json/__init__.py", line 346, in loads
    return _default_decoder.decode(s)
  File "/usr/lib/python3.9/json/decoder.py", line 337, in decode
    obj, end = self.raw_decode(s, idx=_w(s, 0).end())
  File "/usr/lib/python3.9/json/decoder.py", line 355, in raw_decode
    raise JSONDecodeError("Expecting value", s, err.value) from None
json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
故障根因
  • 报错Expecting value: line 1 column 1 (char 0)的直接原因是传入json.loads()的参数是空字符串,解析器从第一个字符开始就没找到任何有效JSON内容,和JSON格式错误无关。
  • 核心差异来自两个运行环境的网络身份不同:本地Thonny发起请求时用的是个人家庭宽带的出口IP,目标电商站点将其识别为普通用户访问,正常返回完整商品页HTML,代码可以正常定位到ld+json脚本标签、拿到有效结构化数据;Trinket作为公共在线编程平台,出口IP属于云服务商公开IP段,早就被目标站点的反爬系统标记为爬虫流量,返回的要么是空白页、要么是人机验证页、要么是缺失结构化数据的异常页面,导致代码要么找不到目标script标签,要么取到的标签文本为空。
  • 代码本身没有做容错处理:默认find()方法一定能找到目标标签、标签内一定有有效内容,一旦请求被拦截就会直接触发解析报错。
修复方案
  • 首先给requests请求添加真实浏览器的User-Agent请求头,模拟普通用户访问,降低被反爬识别的概率。
  • 增加空值校验逻辑,不要默认页面元素一定存在,出现拦截时直接抛出明确提示,避免无意义的JSON解析报错。
  • 调整后的可运行代码示例:
import requests
from bs4 import BeautifulSoup
import json

# 模拟桌面端Chrome浏览器的请求头
REQUEST_HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8"
}

def compare_prices(product_laughs,product_glomark):
    # 发起请求时携带请求头
    resp_lau = requests.get(product_laughs, headers=REQUEST_HEADERS, timeout=10)
    resp_lau.raise_for_status() # 遇到HTTP错误状态码直接抛出异常
    html_lau = resp_lau.content

    resp_glo = requests.get(product_glomark, headers=REQUEST_HEADERS, timeout=10)
    resp_glo.raise_for_status()
    html_glo = resp_glo.content

    soup_lau=BeautifulSoup(html_lau,'html.parser')
    soup_glo=BeautifulSoup(html_glo,'html.parser')
    
    glo_products=soup_glo.find("script", type="application/ld+json")
    # 校验是否找到目标标签
    if not glo_products:
        raise RuntimeError("当前网络环境被目标站点反爬拦截,页面中未找到商品结构化数据标签")
    glomark_content = glo_products.text.strip() # 去除首尾空白、换行字符
    # 校验标签内容是否为空
    if not glomark_content:
        raise RuntimeError("获取到的结构化数据标签内容为空,请求可能被反爬拦截")
    productList=json.loads(glomark_content)
  • 如果添加请求头后在Trinket平台依然报错,说明Trinket的出口IP已经被目标站点完全封禁,公共在线平台的IP池被大量用户共享用来爬取数据,基本没有稳定绕过反爬的可能,这类场景直接切换到本地环境运行即可。

内容的提问来源于stack exchange,提问作者A_Ruwantha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 05:09:18