无法从otto.de抓取全部商品名称,需实现多页商品名爬取
解决Otto.de男士夹克多页商品名称爬取问题
当前代码仅能输出单个商品名称,核心问题有两个:
- 错误定位了商品列表容器而非单个商品元素,导致循环仅执行一次
- 缺少分页逻辑,仅爬取了第一页内容
修改后的完整代码
import time from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 设置要爬取的目标页数 TARGET_PAGES = 3 all_jacket_titles = [] # 配置Chrome浏览器,规避反爬检测 options = webdriver.ChromeOptions() options.add_experimental_option("excludeSwitches", ['enable-automation']) options.add_argument('--disable-blink-features=AutomationControlled') options.add_argument( "User-Agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/110.0.0.0 Safari/537.36") driver = webdriver.Chrome('D:/chromedriver_win32/chromedriver.exe', options=options) driver.get("https://www.otto.de/herren/mode/jacken/") for page_num in range(TARGET_PAGES): print(f"正在爬取第 {page_num + 1} 页...") # 显式等待商品列表加载完成,替代固定sleep更稳定 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "find_tile")) ) page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 定位当前页所有单个商品容器 products = soup.find_all("article", class_="find_tile") for product in products: # 提取商品名称,处理可能的空值情况 title_tag = product.find("a", class_="find_tile__content") if title_tag: title = title_tag.get_text(strip=True) all_jacket_titles.append(title) print(title) # 未到最后一页时,点击下一页按钮 if page_num < TARGET_PAGES - 1: try: next_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, "button[data-test='pagination-next']")) ) next_button.click() time.sleep(2) # 等待页面跳转加载 except Exception as e: print(f"无法继续分页,已完成 {page_num + 1} 页爬取: {str(e)}") break # 关闭浏览器 driver.quit() # 输出所有爬取结果 print("\n全部商品名称:") for idx, title in enumerate(all_jacket_titles, 1): print(f"{idx}. {title}") # 可选:保存结果到CSV文件 # import pandas as pd # pd.DataFrame({"商品名称": all_jacket_titles}).to_csv("otto_男士夹克列表.csv", index=False, encoding="utf-8-sig")
关键修改点说明
- 商品定位修正:将原代码中定位整个列表容器的逻辑,改为定位单个商品的
article.find_tile元素,确保循环遍历当前页所有商品 - 分页逻辑实现:通过循环指定页数,每次爬完当前页后点击下一页按钮,同时处理按钮找不到的异常情况
- 等待优化:用Selenium的显式等待替代固定
time.sleep,确保页面元素加载完成后再解析,提升稳定性 - 结果收集:将所有商品名称存入列表,方便后续查看或保存
内容的提问来源于stack exchange,提问作者Muhammad Umer
相关产品推荐
相关产品推荐

