You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python爬取全分页商品?爬取结果不全问题排查

Newegg.ca 商品爬虫修复方案

问题根源

爬虫出现空结果或抓取不全的核心原因:

  • 分页请求未携带请求头,被网站反爬机制拦截
  • 商品名称匹配依赖文本正则,漏抓标题格式不符的商品
  • 页面元素选择器依赖易变的class组合,稳定性差
  • 异常处理过于宽泛,掩盖实际错误

修复步骤

  1. 全局携带请求头:所有请求传入headers参数,避免被识别为爬虫
  2. 直接定位标题元素:放弃正则文本匹配,通过商品容器的标题标签抓取名称,确保不遗漏
  3. 优化分页解析:用更可靠的方式提取总页数,避免字符串分割出错
  4. 缩小异常捕获范围:只捕获必要异常,方便排查问题
  5. 使用稳定选择器:简化商品容器和价格元素的选择器,降低页面更新导致的失效概率

修复后完整代码

from bs4 import BeautifulSoup
import requests
import re

search_term = input("What product do you want to search for? ")

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
}

# 初始页面获取总页数
url = f"https://www.newegg.ca/p/pl?d={search_term}&N=4131"
response = requests.get(url, headers=headers)
response.raise_for_status()  # 主动抛出请求错误
doc = BeautifulSoup(response.text, "html.parser")

# 更可靠的分页解析
pagination_text = doc.select_one(".list-tool-pagination-text strong")
if not pagination_text:
    print("无法获取分页信息,默认抓取1页")
    pages = 1
else:
    pages = int(pagination_text.text.split("/")[-1].strip())

items_found = {}

for page_num in range(1, pages + 1):
    url = f"https://www.newegg.ca/p/pl?d={search_term}&N=4131&page={page_num}"
    response = requests.get(url, headers=headers)
    response.raise_for_status()
    doc = BeautifulSoup(response.text, "html.parser")

    # 选择所有商品容器
    item_containers = doc.select(".item-container")
    if not item_containers:
        print(f"第{page_num}页未找到商品")
        continue

    for container in item_containers:
        # 抓取商品标题
        title_tag = container.select_one(".item-title")
        if not title_tag:
            continue
        item_name = title_tag.get_text(strip=True)
        # 可选:只保留包含搜索词的商品(不区分大小写)
        if not re.search(search_term, item_name, re.IGNORECASE):
            continue
        
        # 抓取商品链接
        item_link = title_tag['href']
        
        # 抓取价格
        price_strong = container.select_one(".price-current strong")
        if not price_strong:
            continue
        price_str = price_strong.get_text(strip=True).replace(",", "")
        try:
            price = int(price_str)
        except ValueError:
            print(f"商品{item_name}价格解析失败")
            continue
        
        items_found[item_name] = {"price": price, "link": item_link}

# 按价格排序并输出
sorted_items = sorted(items_found.items(), key=lambda x: x[1]['price'])

for item in sorted_items:
    print(item[0])
    print(f"${item[1]['price']}")
    print(item[1]['link'])
    print("-------------------------------")

关键改进说明

  • 请求头更新:使用最新Chrome UA,避免被识别为过时爬虫
  • 元素选择器优化:用select_one/select基于稳定class定位元素,降低失效概率
  • 分页容错:增加分页信息缺失的处理逻辑,避免程序崩溃
  • 价格解析容错:单独捕获价格转换异常,不影响其他商品抓取
  • 标题匹配优化:不区分大小写的正则匹配,避免漏抓带前缀的商品标题

内容的提问来源于stack exchange,提问作者acefz1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 22:05:01