如何用BeautifulSoup实现eBay自定义搜索结果的分页爬取?
解决方案:循环遍历eBay搜索结果的所有分页
你已经搞定了单页数据提取和下一页链接获取,接下来只需要用循环结构重复执行「爬取页面→提取数据→获取下一页链接」的流程,直到没有下一页为止。我帮你修改了代码,直接就能用:
修改后的完整代码
from urllib.request import urlopen as Req from bs4 import BeautifulSoup as souce import time # 用于添加请求延迟,避免被反爬 def scrape_single_page(url, file_handle): """封装单页爬取逻辑,返回下一页链接(没有则返回None)""" try: # 发送请求获取页面内容 client = Req(url) page_html = client.read() client.close() time.sleep(2) # 加2秒延迟,降低被eBay限制的风险 except Exception as e: print(f"请求页面出错啦: {e}") return None # 解析页面 page_soup = souce(page_html, "html.parser") containers = page_soup.findAll("li", {"class":"sresult lvresult clearfix li"}) # 提取当前页商品数据并写入文件 for container in containers: # 处理标题里的逗号(避免CSV列错位),用双引号包裹更稳妥 item_title = container.h3.text.strip().replace('"', "'") item_link = container.h3.a["href"].strip() item_price = container.span.text.strip().replace(",", ".") # 按CSV格式写入,标题用双引号包裹 file_handle.write(f'"{item_title}",{item_link},{item_price}\n') # 获取下一页链接 next_page_td = page_soup.find("td", {"class":"pagn-next"}) if next_page_td and next_page_td.a: return next_page_td.a["href"] else: print("没有更多页面了") return None # 主程序逻辑 start_url = 'https://www.ebay.de/sch/i.html?_fosrp=1&_from=R40&_nkw=iphone&_in_kw=1&_ex_kw=&_sacat=0&_mPrRngCbx=1&_udlo=600&_udhi=4.800&LH_BIN=1&LH_ItemCondition=4&_ftrt=901&_ftrv=1&_sabdlo=&_sabdhi=&_samilow=&_samihi=&_sadis=10&_fpos=&LH_SubLocation=1&_sargn=-1%26saslc%3D0&_fsradio2=%26LH_LocatedIn%3D1&_salic=77&_saact=77&LH_SALE_CURRENCY=0&_sop=2&_dmd=1&_ipg=200' filename = "scrape_ebay.csv" # 用with语句管理文件,自动处理关闭操作 with open(filename, "w", encoding="utf-8") as f: # 写入表头 f.write('item_title,item_link,item_price\n') current_url = start_url # 循环爬取每一页,直到没有下一页 while current_url: print(f"正在爬取页面: {current_url}") current_url = scrape_single_page(current_url, f) print("所有页面爬取完成!")
关键优化点说明:
- 模块化封装:把单页爬取逻辑写成函数,代码更整洁,也方便后续维护。
- 循环控制:用
while循环持续爬取,直到找不到下一页链接时自动终止。 - CSV格式兼容:用双引号包裹商品标题,避免标题里的逗号导致CSV列错位;同时替换了标题里的双引号为单引号,防止格式混乱。
- 反爬防护:添加了
time.sleep(2)的延迟,降低被eBay检测到异常请求的概率。 - 异常处理:简单捕获请求异常,避免单个页面出错导致整个程序崩溃。
- 编码处理:写入文件时指定
utf-8编码,避免特殊字符乱码。
额外提醒:
如果后续发现eBay页面结构更新,比如商品容器的类名、下一页按钮的选择器变化,记得及时调整代码里的findAll和find参数哦。
内容的提问来源于stack exchange,提问作者kir.pir
相关产品推荐
相关产品推荐

