You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多进程爬取Tiffany商品价格遇json.decoder.JSONDecodeError求助

解决Tiffany官网多进程爬取JSON解码错误问题

问题背景

开发了Python脚本爬取Tiffany英国官网项链价格,用requests、BeautifulSoup解析页面,通过multiprocessing.Pool实现多进程批量处理100+商品URL。添加更多URL后出现JSON解码错误,导致爬取中断,怀疑存在失效链接,但不熟悉多进程环境下的错误处理方式。

原脚本代码:

import json
from bs4 import BeautifulSoup
import requests
from multiprocessing import Pool
import pandas as pd

data = {'url':[],'offers_price':[]}

def get_price(url):
    soup = BeautifulSoup(requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}).content, "html.parser")
    data = json.loads(soup.find_all('script', {'type': 'application/ld+json'})[-1].get_text())
    return url, int(data['offers']['price'])

if __name__ == '__main__':

    urls = [
        'https://www.tiffany.co.uk/jewelry/necklaces-pendants/tiffany-infinity-pendant-37725951/',
        'https://www.tiffany.co.uk/jewelry/necklaces-pendants/return-to-tiffany-tiffany-blue-heart-tag-charm-GRP09431/',
        ...
        'https://www.tiffany.co.uk/jewelry/necklaces-pendants/elsa-peretti-diamonds-by-the-yard-sprinkle-necklace-62356812/'
    ]

    with Pool(processes=4) as pool:
            for url, price in pool.imap_unordered(get_price, urls):
                    data['offers_price'].append(price)
                    data['url'].append(url)
    tiffany_necklace = pd.DataFrame(data)

报错信息:

Traceback (most recent call last):
  File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/multiprocessing/pool.py", line 125, in worker
    result = (True, func(*args, **kwds))
  File "/Users/marco/PycharmProjects/hardluxuryprices/tiffany.py", line 11, in get_price
    data = json.loads(soup.find_all('script', {'type': 'application/ld+json'})[-1].get_text())
  File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/json/__init__.py", line 346, in loads
    return _default_decoder.decode(s)
  File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/json/decoder.py", line 337, in decode
    obj, end = self.raw_decode(s, idx=_w(s, 0).end())
  File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/json/decoder.py", line 353, in raw_decode
    obj, end = self.scan_once(s, idx)
json.decoder.JSONDecodeError: Invalid control character at: line 5 column 193 (char 292)

连锁报错:

Traceback (most recent call last):
  File "/Users/marco/PycharmProjects/hardluxuryprices/tiffany.py", line 301, in <module>
    for url, price in pool.imap_unordered(get_price, urls):
  File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/multiprocessing/pool.py", line 870, in next
    raise value
json.decoder.JSONDecodeError: Invalid control character at: line 5 column 193 (char 292)

Process finished with exit code 1

原因分析

  • 核心问题不是链接失效,而是部分页面返回的JSON-LD数据包含无效控制字符(如换行、制表符等不符合JSON规范的字符),导致json.loads解析失败。
  • 多进程环境下,子进程抛出的异常会直接传递到主进程,中断整个imap_unordered迭代,后续有效URL无法继续处理。

解决方法

1. 给get_price函数添加异常捕获

单个URL解析失败时不中断整体爬取,返回错误标识并记录详情,方便后续排查。

2. 预处理JSON字符串,移除无效控制字符

用正则表达式过滤JSON中不允许的控制字符(保留\t、\n、\r)。

3. 给请求添加超时机制

避免请求长时间卡住,影响多进程运行效率。

修改后的完整代码

import json
import re
from bs4 import BeautifulSoup
import requests
from multiprocessing import Pool
import pandas as pd

data = {'url': [], 'offers_price': [], 'status': []}  # 新增status字段记录爬取状态

def clean_json_text(text):
    # 移除JSON不允许的控制字符(保留\t\n\r)
    return re.sub(r'[\x00-\x08\x0b\x0c\x0e-\x1f]', '', text)

def get_price(url):
    try:
        # 添加超时,避免请求卡住
        resp = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}, timeout=10)
        resp.raise_for_status()  # 捕获HTTP错误(比如404、500)
        
        soup = BeautifulSoup(resp.content, "html.parser")
        script_tags = soup.find_all('script', {'type': 'application/ld+json'})
        if not script_tags:
            return url, None, '未找到JSON-LD脚本'
        
        json_text = clean_json_text(script_tags[-1].get_text())
        data = json.loads(json_text)
        
        if 'offers' not in data or 'price' not in data['offers']:
            return url, None, 'JSON结构不符合预期'
        
        price = int(data['offers']['price'])
        return url, price, '成功'
    
    except requests.exceptions.RequestException as e:
        return url, None, f'请求错误: {str(e)}'
    except json.JSONDecodeError as e:
        return url, None, f'JSON解码错误: {str(e)}'
    except Exception as e:
        return url, None, f'其他错误: {str(e)}'

if __name__ == '__main__':
    urls = [
        'https://www.tiffany.co.uk/jewelry/necklaces-pendants/tiffany-infinity-pendant-37725951/',
        'https://www.tiffany.co.uk/jewelry/necklaces-pendants/return-to-tiffany-tiffany-blue-heart-tag-charm-GRP09431/',
        # 其他URL...
        'https://www.tiffany.co.uk/jewelry/necklaces-pendants/elsa-peretti-diamonds-by-the-yard-sprinkle-necklace-62356812/'
    ]

    with Pool(processes=4) as pool:
        # 遍历多进程返回的结果
        for url, price, status in pool.imap_unordered(get_price, urls):
            data['url'].append(url)
            data['offers_price'].append(price)
            data['status'].append(status)
    
    # 生成DataFrame
    tiffany_necklace = pd.DataFrame(data)
    # 可以过滤成功的数据,或者保存全部数据用于排查
    print(tiffany_necklace)
    # 可选:保存到CSV
    tiffany_necklace.to_csv('tiffany_necklaces.csv', index=False)

额外说明

  • 新增的status字段可快速定位问题URL,区分是请求失败、JSON结构变更还是解码错误。
  • 多进程环境下,每个子进程独立处理请求和解析,单个进程的错误不会影响其他进程,主进程只需处理返回的错误结果即可。
  • 若爬取量较大,建议将进程数设为CPU核心数,避免触发网站反爬限制。

内容的提问来源于stack exchange,提问作者Seedizens

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 00:40:44