如何用Python爬取全分页商品?爬取结果不全问题排查
Newegg.ca 商品爬虫修复方案
问题根源
爬虫出现空结果或抓取不全的核心原因:
- 分页请求未携带请求头,被网站反爬机制拦截
- 商品名称匹配依赖文本正则,漏抓标题格式不符的商品
- 页面元素选择器依赖易变的class组合,稳定性差
- 异常处理过于宽泛,掩盖实际错误
修复步骤
- 全局携带请求头:所有请求传入
headers参数,避免被识别为爬虫 - 直接定位标题元素:放弃正则文本匹配,通过商品容器的标题标签抓取名称,确保不遗漏
- 优化分页解析:用更可靠的方式提取总页数,避免字符串分割出错
- 缩小异常捕获范围:只捕获必要异常,方便排查问题
- 使用稳定选择器:简化商品容器和价格元素的选择器,降低页面更新导致的失效概率
修复后完整代码
from bs4 import BeautifulSoup import requests import re search_term = input("What product do you want to search for? ") headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', } # 初始页面获取总页数 url = f"https://www.newegg.ca/p/pl?d={search_term}&N=4131" response = requests.get(url, headers=headers) response.raise_for_status() # 主动抛出请求错误 doc = BeautifulSoup(response.text, "html.parser") # 更可靠的分页解析 pagination_text = doc.select_one(".list-tool-pagination-text strong") if not pagination_text: print("无法获取分页信息,默认抓取1页") pages = 1 else: pages = int(pagination_text.text.split("/")[-1].strip()) items_found = {} for page_num in range(1, pages + 1): url = f"https://www.newegg.ca/p/pl?d={search_term}&N=4131&page={page_num}" response = requests.get(url, headers=headers) response.raise_for_status() doc = BeautifulSoup(response.text, "html.parser") # 选择所有商品容器 item_containers = doc.select(".item-container") if not item_containers: print(f"第{page_num}页未找到商品") continue for container in item_containers: # 抓取商品标题 title_tag = container.select_one(".item-title") if not title_tag: continue item_name = title_tag.get_text(strip=True) # 可选:只保留包含搜索词的商品(不区分大小写) if not re.search(search_term, item_name, re.IGNORECASE): continue # 抓取商品链接 item_link = title_tag['href'] # 抓取价格 price_strong = container.select_one(".price-current strong") if not price_strong: continue price_str = price_strong.get_text(strip=True).replace(",", "") try: price = int(price_str) except ValueError: print(f"商品{item_name}价格解析失败") continue items_found[item_name] = {"price": price, "link": item_link} # 按价格排序并输出 sorted_items = sorted(items_found.items(), key=lambda x: x[1]['price']) for item in sorted_items: print(item[0]) print(f"${item[1]['price']}") print(item[1]['link']) print("-------------------------------")
关键改进说明
- 请求头更新:使用最新Chrome UA,避免被识别为过时爬虫
- 元素选择器优化:用
select_one/select基于稳定class定位元素,降低失效概率 - 分页容错:增加分页信息缺失的处理逻辑,避免程序崩溃
- 价格解析容错:单独捕获价格转换异常,不影响其他商品抓取
- 标题匹配优化:不区分大小写的正则匹配,避免漏抓带前缀的商品标题
内容的提问来源于stack exchange,提问作者acefz1
相关产品推荐
相关产品推荐

