You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Windows正常运行的Python脚本在Debian报JSONDecodeError求助

问题

我编写了一个基于BeautifulSoup的Python脚本,在Windows 10上运行完全正常,但部署到Debian 11时出现JSON解析错误,无法定位原因。

脚本代码

from bs4 import BeautifulSoup
import re
import json
import requests
import urllib3
import time
import random

user_agents = [
  "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36 OPR/101.0.0.0",
  "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36"
  ]
random_user_agent = random.choice(user_agents)
headers = {
    'User-Agent': random_user_agent
}

url = "https://www.alza.hu/intel-core-i9-13900ks-d7603014.htm"
session_object = requests.Session()
response = session_object.get(url, headers=headers, timeout=20)
get_source = (response).text
print(response.status_code)
time.sleep(5)
soup = BeautifulSoup(get_source, "lxml")
json_schema = soup.find_all('script', attrs={'type': 'application/ld+json'})[1]
json_file = json.loads(json_schema.get_text())
sku = json_file['sku']
price = json_file['offers']['price']
stock = json_file['offers']['availability']
print(sku, price, stock)    

错误信息

Windows 10上无运行问题,但Debian 11执行时输出状态码200后抛出错误:

Traceback (most recent call last):
json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)

已确认两台系统依赖版本一致,Debian可正常访问目标URL。打印json_schema显示正常,但解析为JSON时触发错误。


排查与解决思路

  • 统一响应编码:Windows和Linux环境下,requests对响应编码的自动识别可能存在差异,导致文本解析出现空白或乱码。强制指定编码可解决该问题:

    response.encoding = 'utf-8'
    get_source = response.text
    

    或直接解码二进制内容:

    get_source = response.content.decode('utf-8')
    
  • 清理JSON文本中的无效字符:即使json_schema打印正常,实际文本可能包含BOM(字节顺序标记)或不可见空白字符,导致JSON解析失败。添加清理步骤:

    json_text = json_schema.get_text().strip()
    # 移除UTF-8 BOM
    if json_text.startswith('\ufeff'):
        json_text = json_text[1:]
    json_file = json.loads(json_text)
    
  • 对比跨环境响应内容:目标网站可能根据访问环境(IP、系统特征)返回不同内容,即使状态码为200。在Debian上保存响应到文件,对比Windows下的内容:

    with open('debian_response.html', 'w', encoding='utf-8') as f:
        f.write(get_source)
    

    查看文件中第二个application/ld+json脚本的实际内容,确认是否为空或格式异常。

  • 避免硬编码索引依赖:find_all(...)[1]依赖固定的脚本顺序,网站可能动态调整脚本位置。改用内容匹配筛选目标JSON:

    target_json = None
    for script in soup.find_all('script', attrs={'type': 'application/ld+json'}):
        try:
            data = json.loads(script.get_text().strip())
            if 'sku' in data and 'offers' in data:
                target_json = data
                break
        except json.JSONDecodeError:
            continue
    if target_json:
        sku = target_json['sku']
        price = target_json['offers']['price']
        stock = target_json['offers']['availability']
        print(sku, price, stock)
    else:
        print("未找到目标JSON数据")
    

内容的提问来源于stack exchange,提问作者Zoltan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 03:24:59