You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium使用for-in-loop无法逐页获取爬取数据的问题求助

问题原因

  • 分页循环参数错误:Python的range()为左闭右开区间,你写的range(1,2)仅会生成1这一个值,循环只会执行1次,自然只能获取到第1页数据。
  • 元素引用失效:获取到列表页的商品链接元素后直接跳转到详情页,原列表页的DOM被覆盖,剩余未遍历的链接元素会直接失效,无法正常取值。
  • 数据存储逻辑错误:存储商品数据的products列表定义在单商品循环内部,每次循环都会被重置,且爬取到的数据没有存入全局的output列表,无法汇总全量爬取结果。

修正后代码

# 依赖安装(已安装可注释)
!pip install selenium
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.wait import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
from bs4 import BeautifulSoup

browser = webdriver.Chrome(executable_path='./chromedriver.exe')
wait = WebDriverWait(browser,5)
output = list()
# 要爬取的总页数,比如爬前10页就改为range(1, 11)
total_page = 5
for i in range(1, total_page+1): 
    browser.get(f"https://www.rakuten.com.tw/shop/watsons/product/?l-id=tw_shop_inshop_cat&p={i}")
    # 等待列表加载完成
    wait.until(EC.presence_of_element_located((By.XPATH,"//div[@class='b-content b-fix-2lines']")))
    # 先把当前页所有商品链接提取为字符串存储,避免后续跳转后元素失效
    product_link_elements = browser.find_elements(By.XPATH,"//div[@class='b-content b-fix-2lines']/b/a")
    current_page_links = [link.get_attribute('href') for link in product_link_elements]
    
    # 逐个爬取商品详情
    for link in current_page_links:
        print(f"正在爬取:{link}")
        browser.get(link)
        # 等待详情页加载,可根据需要调整等待条件
        time.sleep(1)
        soup = BeautifulSoup(browser.page_source, 'lxml')
        try:
            product = {}
            product['商品名稱'] = soup.find('div',class_="b-subarea b-layout-right shop-item ng-scope").h1.text.replace('\n','').strip()
            product['價錢'] = soup.find('strong',class_="b-text-xlarge qa-product-actualPrice").text.replace('\n','').strip()
            all_data = soup.find_all("div",class_="b-container-child")[2]
            main_data = all_data.find_all("span")[-1]
            product['購買次數'] = main_data.text.strip()
            output.append(product)
            print(product)
        except Exception as e:
            print(f"商品{link}爬取失败:{str(e)}")
            continue

# 所有页爬完后可打印汇总结果
print(f"共爬取到{len(output)}条商品数据")
print(output)

注意事项

  • 可根据自己的爬取需求调整total_page的数值,确认爬取的总页数
  • 可以适当增加详情页的等待时间,避免页面未加载完成就解析导致报错
  • 新增了异常捕获逻辑,避免个别商品页结构异常导致整个爬取任务中断

内容的提问来源于stack exchange,提问作者鄭鼎彥

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 18:06:06