多进程爬取Tiffany商品价格遇json.decoder.JSONDecodeError求助
解决Tiffany官网多进程爬取JSON解码错误问题
问题背景
开发了Python脚本爬取Tiffany英国官网项链价格,用requests、BeautifulSoup解析页面,通过multiprocessing.Pool实现多进程批量处理100+商品URL。添加更多URL后出现JSON解码错误,导致爬取中断,怀疑存在失效链接,但不熟悉多进程环境下的错误处理方式。
原脚本代码:
import json from bs4 import BeautifulSoup import requests from multiprocessing import Pool import pandas as pd data = {'url':[],'offers_price':[]} def get_price(url): soup = BeautifulSoup(requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}).content, "html.parser") data = json.loads(soup.find_all('script', {'type': 'application/ld+json'})[-1].get_text()) return url, int(data['offers']['price']) if __name__ == '__main__': urls = [ 'https://www.tiffany.co.uk/jewelry/necklaces-pendants/tiffany-infinity-pendant-37725951/', 'https://www.tiffany.co.uk/jewelry/necklaces-pendants/return-to-tiffany-tiffany-blue-heart-tag-charm-GRP09431/', ... 'https://www.tiffany.co.uk/jewelry/necklaces-pendants/elsa-peretti-diamonds-by-the-yard-sprinkle-necklace-62356812/' ] with Pool(processes=4) as pool: for url, price in pool.imap_unordered(get_price, urls): data['offers_price'].append(price) data['url'].append(url) tiffany_necklace = pd.DataFrame(data)
报错信息:
Traceback (most recent call last): File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/multiprocessing/pool.py", line 125, in worker result = (True, func(*args, **kwds)) File "/Users/marco/PycharmProjects/hardluxuryprices/tiffany.py", line 11, in get_price data = json.loads(soup.find_all('script', {'type': 'application/ld+json'})[-1].get_text()) File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/json/__init__.py", line 346, in loads return _default_decoder.decode(s) File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/json/decoder.py", line 337, in decode obj, end = self.raw_decode(s, idx=_w(s, 0).end()) File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/json/decoder.py", line 353, in raw_decode obj, end = self.scan_once(s, idx) json.decoder.JSONDecodeError: Invalid control character at: line 5 column 193 (char 292)
连锁报错:
Traceback (most recent call last): File "/Users/marco/PycharmProjects/hardluxuryprices/tiffany.py", line 301, in <module> for url, price in pool.imap_unordered(get_price, urls): File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/multiprocessing/pool.py", line 870, in next raise value json.decoder.JSONDecodeError: Invalid control character at: line 5 column 193 (char 292) Process finished with exit code 1
原因分析
- 核心问题不是链接失效,而是部分页面返回的JSON-LD数据包含无效控制字符(如换行、制表符等不符合JSON规范的字符),导致
json.loads解析失败。 - 多进程环境下,子进程抛出的异常会直接传递到主进程,中断整个
imap_unordered迭代,后续有效URL无法继续处理。
解决方法
1. 给get_price函数添加异常捕获
单个URL解析失败时不中断整体爬取,返回错误标识并记录详情,方便后续排查。
2. 预处理JSON字符串,移除无效控制字符
用正则表达式过滤JSON中不允许的控制字符(保留\t、\n、\r)。
3. 给请求添加超时机制
避免请求长时间卡住,影响多进程运行效率。
修改后的完整代码
import json import re from bs4 import BeautifulSoup import requests from multiprocessing import Pool import pandas as pd data = {'url': [], 'offers_price': [], 'status': []} # 新增status字段记录爬取状态 def clean_json_text(text): # 移除JSON不允许的控制字符(保留\t\n\r) return re.sub(r'[\x00-\x08\x0b\x0c\x0e-\x1f]', '', text) def get_price(url): try: # 添加超时,避免请求卡住 resp = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}, timeout=10) resp.raise_for_status() # 捕获HTTP错误(比如404、500) soup = BeautifulSoup(resp.content, "html.parser") script_tags = soup.find_all('script', {'type': 'application/ld+json'}) if not script_tags: return url, None, '未找到JSON-LD脚本' json_text = clean_json_text(script_tags[-1].get_text()) data = json.loads(json_text) if 'offers' not in data or 'price' not in data['offers']: return url, None, 'JSON结构不符合预期' price = int(data['offers']['price']) return url, price, '成功' except requests.exceptions.RequestException as e: return url, None, f'请求错误: {str(e)}' except json.JSONDecodeError as e: return url, None, f'JSON解码错误: {str(e)}' except Exception as e: return url, None, f'其他错误: {str(e)}' if __name__ == '__main__': urls = [ 'https://www.tiffany.co.uk/jewelry/necklaces-pendants/tiffany-infinity-pendant-37725951/', 'https://www.tiffany.co.uk/jewelry/necklaces-pendants/return-to-tiffany-tiffany-blue-heart-tag-charm-GRP09431/', # 其他URL... 'https://www.tiffany.co.uk/jewelry/necklaces-pendants/elsa-peretti-diamonds-by-the-yard-sprinkle-necklace-62356812/' ] with Pool(processes=4) as pool: # 遍历多进程返回的结果 for url, price, status in pool.imap_unordered(get_price, urls): data['url'].append(url) data['offers_price'].append(price) data['status'].append(status) # 生成DataFrame tiffany_necklace = pd.DataFrame(data) # 可以过滤成功的数据,或者保存全部数据用于排查 print(tiffany_necklace) # 可选:保存到CSV tiffany_necklace.to_csv('tiffany_necklaces.csv', index=False)
额外说明
- 新增的
status字段可快速定位问题URL,区分是请求失败、JSON结构变更还是解码错误。 - 多进程环境下,每个子进程独立处理请求和解析,单个进程的错误不会影响其他进程,主进程只需处理返回的错误结果即可。
- 若爬取量较大,建议将进程数设为CPU核心数,避免触发网站反爬限制。
内容的提问来源于stack exchange,提问作者Seedizens
相关产品推荐
相关产品推荐

