You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法从otto.de抓取全部商品名称,需实现多页商品名爬取

解决Otto.de男士夹克多页商品名称爬取问题

当前代码仅能输出单个商品名称,核心问题有两个:

  1. 错误定位了商品列表容器而非单个商品元素,导致循环仅执行一次
  2. 缺少分页逻辑,仅爬取了第一页内容

修改后的完整代码

import time
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 设置要爬取的目标页数
TARGET_PAGES = 3
all_jacket_titles = []

# 配置Chrome浏览器,规避反爬检测
options = webdriver.ChromeOptions()
options.add_experimental_option("excludeSwitches", ['enable-automation'])
options.add_argument('--disable-blink-features=AutomationControlled')
options.add_argument(
    "User-Agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/110.0.0.0 Safari/537.36")

driver = webdriver.Chrome('D:/chromedriver_win32/chromedriver.exe', options=options)
driver.get("https://www.otto.de/herren/mode/jacken/")

for page_num in range(TARGET_PAGES):
    print(f"正在爬取第 {page_num + 1} 页...")
    # 显式等待商品列表加载完成,替代固定sleep更稳定
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "find_tile"))
    )
    
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, 'html.parser')
    
    # 定位当前页所有单个商品容器
    products = soup.find_all("article", class_="find_tile")
    for product in products:
        # 提取商品名称,处理可能的空值情况
        title_tag = product.find("a", class_="find_tile__content")
        if title_tag:
            title = title_tag.get_text(strip=True)
            all_jacket_titles.append(title)
            print(title)
    
    # 未到最后一页时,点击下一页按钮
    if page_num < TARGET_PAGES - 1:
        try:
            next_button = WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.CSS_SELECTOR, "button[data-test='pagination-next']"))
            )
            next_button.click()
            time.sleep(2)  # 等待页面跳转加载
        except Exception as e:
            print(f"无法继续分页,已完成 {page_num + 1} 页爬取: {str(e)}")
            break

# 关闭浏览器
driver.quit()

# 输出所有爬取结果
print("\n全部商品名称:")
for idx, title in enumerate(all_jacket_titles, 1):
    print(f"{idx}. {title}")

# 可选:保存结果到CSV文件
# import pandas as pd
# pd.DataFrame({"商品名称": all_jacket_titles}).to_csv("otto_男士夹克列表.csv", index=False, encoding="utf-8-sig")

关键修改点说明

  • 商品定位修正:将原代码中定位整个列表容器的逻辑,改为定位单个商品的article.find_tile元素,确保循环遍历当前页所有商品
  • 分页逻辑实现:通过循环指定页数,每次爬完当前页后点击下一页按钮,同时处理按钮找不到的异常情况
  • 等待优化:用Selenium的显式等待替代固定time.sleep,确保页面元素加载完成后再解析,提升稳定性
  • 结果收集:将所有商品名称存入列表,方便后续查看或保存

内容的提问来源于stack exchange,提问作者Muhammad Umer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 12:13:09