You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Allen Solly爬取Load More按钮点击报错:元素无法滚动到视图

解决Allen Solly网站爬取时Load More按钮无法点击的ElementNotInteractableException问题

我尝试爬取Allen Solly网站的100件男士衬衫数据,页面初始仅展示30-32件商品,因此添加了点击“Load More”按钮加载更多的逻辑,但点击按钮时触发selenium.common.exceptions.ElementNotInteractableException,错误提示“Element could not be scrolled into view”。

错误栈信息

selenium.common.exceptions.ElementNotInteractableException: Message: Element

问题分析

出现该错误的核心原因:

  • 原代码仅滚动到页面底部,但Load More按钮可能未完全进入可视区域,或被固定导航栏等元素遮挡
  • 循环逻辑存在重复滚动和点击的问题,导致按钮状态不稳定
  • 直接调用click()方法对动态渲染的元素兼容性较差

修复方案

针对上述问题,做以下关键调整:

  • 精准滚动到按钮位置:定位按钮后将其滚动到视图中央,确保完全可见
  • JavaScript点击替代原生click:绕过元素交互状态的限制
  • 优化循环逻辑:每次点击后等待新商品加载,直到数量达标或按钮消失
  • 产品去重:避免重复添加已爬取的商品

修改后的完整代码

import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

# 初始化Firefox驱动
driver = webdriver.Firefox()

try:
    # 访问男士衬衫页面
    url = 'https://allensolly.abfrl.in/c/men-shirts'
    driver.get(url)

    # 等待页面初始商品加载完成
    WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CLASS_NAME, 'ProductCard_productInfo__uZhFN')))

    product_info = []
    seen_titles = set()  # 用于商品去重

    while len(product_info) < 100:
        # 提取当前页面所有商品并去重
        page_source = driver.page_source
        soup = BeautifulSoup(page_source, 'html.parser')
        products = soup.find_all('div', class_='ProductCard_productInfo__uZhFN')
        
        for product in products:
            title_element = product.find('div', class_='ProductCard_title__9M6wy')
            detail_element = product.find('div', class_='ProductCard_description__BQzle')
            if title_element and detail_element:
                title = title_element.text.strip()
                detail = detail_element.text.strip()
                if title not in seen_titles:
                    seen_titles.add(title)
                    product_info.append((title, detail))
        
        # 商品数量达标则停止循环
        if len(product_info) >= 100:
            break

        # 尝试定位并点击Load More按钮
        try:
            # 等待按钮出现
            load_more_button = WebDriverWait(driver, 10).until(
                EC.presence_of_element_located((By.XPATH, '//button[contains(text(), "LOAD MORE")]'))
            )
            # 将按钮滚动到视图中央,避免被遮挡
            driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", load_more_button)
            time.sleep(1)  # 等待滚动动画完成
            # 用JavaScript点击按钮,绕过交互限制
            driver.execute_script("arguments[0].click();", load_more_button)
            print("点击了'LOAD MORE'按钮")
            
            # 等待新商品加载完成
            WebDriverWait(driver, 10).until(
                lambda d: len(d.find_elements(By.CLASS_NAME, 'ProductCard_productInfo__uZhFN')) > len(products)
            )
            time.sleep(2)  # 额外等待确保渲染完成
        except Exception as e:
            print(f"无更多商品或按钮无法点击: {str(e)}")
            break

    print(f"共获取到{len(product_info)}件唯一商品")

    # 爬取每个商品的详情信息
    data = []
    for idx, (title, detail) in enumerate(product_info[:100], start=1):
        print(f"{idx}. 标题: {title}\n   详情: {detail}\n")
        try:
            # 定位目标商品元素
            product_element = WebDriverWait(driver, 10).until(
                EC.presence_of_element_located((By.XPATH, f'//div[contains(@class, "ProductCard_title__9M6wy") and text()="{title}"]/ancestor::div[contains(@class, "ProductCard_productInfo__uZhFN")]'))
            )
            # 滚动到商品位置并点击打开详情
            driver.execute_script("arguments[0].scrollIntoView(true);", product_element)
            driver.execute_script("arguments[0].click();", product_element)
            print("打开商品详情页")
            
            # 等待新标签页打开并切换
            WebDriverWait(driver, 10).until(lambda d: len(d.window_handles) > 1)
            driver.switch_to.window(driver.window_handles[-1])
            
            # 爬取商品描述
            product_description = WebDriverWait(driver, 10).until(
                EC.presence_of_element_located((By.CSS_SELECTOR, '.ProductDetails_description__7hqm9'))
            ).text.strip()
            
            data.append(('Allen Solly', 'Men', title, detail, product_description))
            print(f"商品描述: {product_description[:100]}...")  # 仅展示前100字符
            
            # 关闭详情页并切回商品列表页
            driver.close()
            driver.switch_to.window(driver.window_handles[0])
            # 等待列表页恢复可交互状态
            WebDriverWait(driver, 10).until(
                EC.presence_of_element_located((By.CLASS_NAME, 'ProductCard_productInfo__uZhFN'))
            )
        except Exception as e:
            print(f"处理商品详情时出错: {str(e)}")
            # 出错后确保切回列表页
            if len(driver.window_handles) > 1:
                driver.close()
                driver.switch_to.window(driver.window_handles[0])

    # 将数据保存到Excel文件
    df = pd.DataFrame(data, columns=['品牌', '分类', '标题', '商品详情', '商品描述'])
    df.to_excel('allen_solly_shirts.xlsx', index=False)
    print("数据已保存到allen_solly_shirts.xlsx")

finally:
    driver.quit()

内容的提问来源于stack exchange,提问作者Vinay Sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 05:28:13