Windows正常运行的Python脚本在Debian报JSONDecodeError求助
问题
我编写了一个基于BeautifulSoup的Python脚本,在Windows 10上运行完全正常,但部署到Debian 11时出现JSON解析错误,无法定位原因。
脚本代码
from bs4 import BeautifulSoup import re import json import requests import urllib3 import time import random user_agents = [ "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36 OPR/101.0.0.0", "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36" ] random_user_agent = random.choice(user_agents) headers = { 'User-Agent': random_user_agent } url = "https://www.alza.hu/intel-core-i9-13900ks-d7603014.htm" session_object = requests.Session() response = session_object.get(url, headers=headers, timeout=20) get_source = (response).text print(response.status_code) time.sleep(5) soup = BeautifulSoup(get_source, "lxml") json_schema = soup.find_all('script', attrs={'type': 'application/ld+json'})[1] json_file = json.loads(json_schema.get_text()) sku = json_file['sku'] price = json_file['offers']['price'] stock = json_file['offers']['availability'] print(sku, price, stock)
错误信息
Windows 10上无运行问题,但Debian 11执行时输出状态码200后抛出错误:
Traceback (most recent call last): json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
已确认两台系统依赖版本一致,Debian可正常访问目标URL。打印json_schema显示正常,但解析为JSON时触发错误。
排查与解决思路
统一响应编码:Windows和Linux环境下,requests对响应编码的自动识别可能存在差异,导致文本解析出现空白或乱码。强制指定编码可解决该问题:
response.encoding = 'utf-8' get_source = response.text或直接解码二进制内容:
get_source = response.content.decode('utf-8')清理JSON文本中的无效字符:即使
json_schema打印正常,实际文本可能包含BOM(字节顺序标记)或不可见空白字符,导致JSON解析失败。添加清理步骤:json_text = json_schema.get_text().strip() # 移除UTF-8 BOM if json_text.startswith('\ufeff'): json_text = json_text[1:] json_file = json.loads(json_text)对比跨环境响应内容:目标网站可能根据访问环境(IP、系统特征)返回不同内容,即使状态码为200。在Debian上保存响应到文件,对比Windows下的内容:
with open('debian_response.html', 'w', encoding='utf-8') as f: f.write(get_source)查看文件中第二个
application/ld+json脚本的实际内容,确认是否为空或格式异常。避免硬编码索引依赖:
find_all(...)[1]依赖固定的脚本顺序,网站可能动态调整脚本位置。改用内容匹配筛选目标JSON:target_json = None for script in soup.find_all('script', attrs={'type': 'application/ld+json'}): try: data = json.loads(script.get_text().strip()) if 'sku' in data and 'offers' in data: target_json = data break except json.JSONDecodeError: continue if target_json: sku = target_json['sku'] price = target_json['offers']['price'] stock = target_json['offers']['availability'] print(sku, price, stock) else: print("未找到目标JSON数据")
内容的提问来源于stack exchange,提问作者Zoltan
相关产品推荐
相关产品推荐

