Gumtree多页面爬取脚本无输出问题求助及修复需求
修复Gumtree法拉利商品多页面爬虫脚本
问题根源
- 反爬拦截:Gumtree会拦截无请求头的请求,直接返回空页面,导致脚本无输出
- 选择器失效:原脚本使用的
h3-responsive类名已不是当前页面中商品名称、价格的对应标签类 - 分页选择器错误:原分页定位规则不符合当前Gumtree的分页结构
- 变量名冲突:循环内变量与全局变量重名,易引发逻辑混淆
修复后的完整脚本
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin # 初始搜索URL base_url = "https://www.gumtree.com/search?search_category=all&q=ferrari" # 模拟浏览器请求头,规避反爬 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } current_url = base_url while current_url: # 发送带请求头的HTTP请求 response = requests.get(current_url, headers=headers) # 检查请求是否成功,失败则抛出异常 response.raise_for_status() soup = BeautifulSoup(response.text, "html.parser") # 获取所有商品容器 listings = soup.find_all("div", class_="listing-item-container") for listing in listings: # 提取商品名称,处理标签不存在的情况 name_element = listing.find("a", class_="listing-title") item_name = name_element.text.strip() if name_element else "未获取到名称" # 提取商品价格,处理标签不存在的情况 price_element = listing.find("span", class_="listing-price") item_price = price_element.text.strip() if price_element else "未获取到价格" print(f"名称: {item_name} | 价格: {item_price}") # 定位下一页链接 next_page_link = soup.select_one("li.next>a") if next_page_link: next_url = next_page_link.get("href") current_url = urljoin(base_url, next_url) print(f"\n===== 爬取下一页: {current_url} =====") else: current_url = None print("\n===== 所有页面爬取完成 =====")
关键修改说明
- 添加请求头:通过
User-Agent模拟浏览器访问,避免被Gumtree的反爬机制拦截 - 更新选择器:根据当前Gumtree页面结构,使用
listing-title(商品名)、listing-price(价格)、li.next>a(下一页)的选择规则 - 容错处理:对可能不存在的标签做判断,避免脚本中途报错终止
- 结构优化:先获取商品容器,再在容器内提取信息,提升数据提取的准确性
- 错误排查:加入
response.raise_for_status(),请求失败时直接抛出异常,便于定位问题
内容的提问来源于stack exchange,提问作者Jack9992
相关产品推荐
相关产品推荐

