You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy无法获取页面隐藏数据:加载更多按钮响应问题求助

Fixing Hidden Content Scraping in TopAchat Without Selenium

Hey there! Let's break down what's going on here and fix your scraping issue without jumping to Selenium just yet.

Why Your Current Approach Isn't Working

First, you're right that the "voir plus de résultats" content is already loaded but hidden—no AJAX means the data is sitting in the raw HTML, just hidden by CSS. Your mistake is trying to use scrapy.FormRequest.from_response to "click" a href="#" element:

  • That anchor link is just a frontend trigger for a JavaScript show/hide action, not a form submission. Scrapy's FormRequest is designed for submitting form data, not simulating JS clicks.
  • The error no clickable element matching clickdata {'href' :"#"} happens because Scrapy can't find a form-associated element with that href to trigger.

The Fix: Extract Hidden Content Directly

Since the content is already in the page's HTML, you don't need to "click" anything—just target the hidden elements in your XPath selectors. Here's how to adjust your code:

Step 1: Update Your Crawl Rules

Check the page's raw HTML (right-click > "View Page Source" in your browser) to find the hidden product containers. They might have classes like hidden, d-none, or another class that hides them via CSS. Modify your LinkExtractor to include these elements:

import scrapy
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
from TopAchat.items import TopachatItem

class TopAchatSpider(CrawlSpider):
    name = "TopAchat1"
    # Directly target the Apple page instead of the marque list to save a crawl step
    start_urls = ["https://www.topachat.com/pages/lexique2_marque_est_APPLE_puis_cat_est_accueil.html"]  

    rules = (
        # Include both visible AND hidden product links
        Rule(LinkExtractor(restrict_xpaths=(
            '//*[@id="content"]/section/ul[@class="listProduct"]//a',
            # Adjust this XPath to match the actual hidden container class in the page source
            '//*[@id="content"]/section/ul[contains(@class, "hidden")]//a'  
        )), callback='parse_item'),
    )

    def parse_item(self, response):
        item = TopachatItem()
        item["brandname"] = response.xpath("//div[@class='small']/div/text()").extract_first()[4:].split(" ")[0]
        item["href"] = response.request.url
        item["name"] = response.xpath("//div[@class='libelle']/h1/text()").extract_first()
        item["specs"] = response.xpath("//strong[@class='big short-descr']/text()").extract_first()
        item["product_type"] = response.xpath("//nav[@class='meta small']//p//b/text()").extract_first()
        
        yield item  # Yield the item properly—no need for the FormRequest anymore

Step 2: Verify the Hidden Container Selector

Make sure the XPath for the hidden list matches what you see in the page source. For example, if the hidden products are in a ul with class listProduct hidden, adjust the XPath to '//ul[@class="listProduct hidden"]//a'.

Key Notes

  • Scrapy sees raw HTML: Unlike browsers, Scrapy doesn't care about CSS visibility—it reads the full HTML sent by the server, so hidden elements are just as accessible as visible ones.
  • Ditch the FormRequest: You don't need it here because there's no form to submit. The "voir plus" button only toggles CSS visibility, not data loading.
  • Start directly on the target page: I updated your start_urls to point straight to the Apple page instead of the marque list—this saves you a crawl step if that's your target.

If you still can't find the hidden elements, double-check the page source (not the dev tools' Elements tab, which shows rendered content) to confirm where the extra products are stored. Sometimes they're in a separate container with a different ID or class.

内容的提问来源于stack exchange,提问作者AnasElM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:29:17