使用Python Selenium实现多页循环爬取并自动翻页
多页面数据爬取:从单页循环到自动翻页的解决方案
我原本用for循环爬取列表数据,但目标数据分散在多个页面,需要实现爬完当前页后自动点击下一页,重复爬取直到最后一页。
初始单页爬取代码
base=[] for product in products: product.click() info=driver.find_element(By.Xpath,'asdfsd').text time.sleep(1) list={'Info':info} base.append(list) driver.execute_script("window.history.go(-1)") time.sleep(2)
多页自动爬取解决方案
通过while True循环实现自动翻页,同时补充异常处理来终止循环(原代码缺少异常捕获部分,会导致最后一页报错):
from selenium.common.exceptions import NoSuchElementException base=[] while True: products=driver.find_elements(By.XPATH,'asdfs') for product in products: product.click() info=driver.find_element(By.Xpath,'asdfsd').text time.sleep(1) item={'Info':info} # 避免用list作为变量名,和内置类型冲突 base.append(item) driver.execute_script("window.history.go(-1)") time.sleep(2) try: nextpage=driver.find_element(By.XPATH,'assfasd') nextpage.click() time.sleep(2) except NoSuchElementException: # 找不到下一页按钮,说明已到最后一页,终止循环 break
关键说明
- 用
while True创建无限循环,每次循环处理当前页面的所有数据 - 内层for循环负责爬取当前页的每个条目
- 外层try-except块用于检测下一页按钮:找到则点击进入下一页,找不到则跳出循环,结束爬取
- 额外优化:把变量名
list改成item,避免和Python内置的list类型冲突
内容的提问来源于stack exchange,提问作者SERGIO
相关产品推荐
相关产品推荐

