You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup实现eBay自定义搜索结果的分页爬取?

解决方案:循环遍历eBay搜索结果的所有分页

你已经搞定了单页数据提取和下一页链接获取,接下来只需要用循环结构重复执行「爬取页面→提取数据→获取下一页链接」的流程,直到没有下一页为止。我帮你修改了代码,直接就能用:

修改后的完整代码

from urllib.request import urlopen as Req
from bs4 import BeautifulSoup as souce
import time  # 用于添加请求延迟,避免被反爬

def scrape_single_page(url, file_handle):
    """封装单页爬取逻辑,返回下一页链接(没有则返回None)"""
    try:
        # 发送请求获取页面内容
        client = Req(url)
        page_html = client.read()
        client.close()
        time.sleep(2)  # 加2秒延迟,降低被eBay限制的风险
    except Exception as e:
        print(f"请求页面出错啦: {e}")
        return None
    
    # 解析页面
    page_soup = souce(page_html, "html.parser")
    containers = page_soup.findAll("li", {"class":"sresult lvresult clearfix li"})
    
    # 提取当前页商品数据并写入文件
    for container in containers:
        # 处理标题里的逗号(避免CSV列错位),用双引号包裹更稳妥
        item_title = container.h3.text.strip().replace('"', "'")
        item_link = container.h3.a["href"].strip()
        item_price = container.span.text.strip().replace(",", ".")
        
        # 按CSV格式写入,标题用双引号包裹
        file_handle.write(f'"{item_title}",{item_link},{item_price}\n')
    
    # 获取下一页链接
    next_page_td = page_soup.find("td", {"class":"pagn-next"})
    if next_page_td and next_page_td.a:
        return next_page_td.a["href"]
    else:
        print("没有更多页面了")
        return None

# 主程序逻辑
start_url = 'https://www.ebay.de/sch/i.html?_fosrp=1&_from=R40&_nkw=iphone&_in_kw=1&_ex_kw=&_sacat=0&_mPrRngCbx=1&_udlo=600&_udhi=4.800&LH_BIN=1&LH_ItemCondition=4&_ftrt=901&_ftrv=1&_sabdlo=&_sabdhi=&_samilow=&_samihi=&_sadis=10&_fpos=&LH_SubLocation=1&_sargn=-1%26saslc%3D0&_fsradio2=%26LH_LocatedIn%3D1&_salic=77&_saact=77&LH_SALE_CURRENCY=0&_sop=2&_dmd=1&_ipg=200'
filename = "scrape_ebay.csv"

# 用with语句管理文件,自动处理关闭操作
with open(filename, "w", encoding="utf-8") as f:
    # 写入表头
    f.write('item_title,item_link,item_price\n')
    
    current_url = start_url
    # 循环爬取每一页,直到没有下一页
    while current_url:
        print(f"正在爬取页面: {current_url}")
        current_url = scrape_single_page(current_url, f)

print("所有页面爬取完成!")

关键优化点说明:

  • 模块化封装:把单页爬取逻辑写成函数,代码更整洁,也方便后续维护。
  • 循环控制:用while循环持续爬取,直到找不到下一页链接时自动终止。
  • CSV格式兼容:用双引号包裹商品标题,避免标题里的逗号导致CSV列错位;同时替换了标题里的双引号为单引号,防止格式混乱。
  • 反爬防护:添加了time.sleep(2)的延迟,降低被eBay检测到异常请求的概率。
  • 异常处理:简单捕获请求异常,避免单个页面出错导致整个程序崩溃。
  • 编码处理:写入文件时指定utf-8编码,避免特殊字符乱码。

额外提醒:

如果后续发现eBay页面结构更新,比如商品容器的类名、下一页按钮的选择器变化,记得及时调整代码里的findAll和find参数哦。

内容的提问来源于stack exchange,提问作者kir.pir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:56:17