You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python Scrapy中提取JS onclick属性的href列表并循环解析

Got it, let's walk through exactly how to extract those JavaScript-driven href values from your pagination HTML using Scrapy, and set up a loop to parse each page.

First, here's the pagination HTML we're working with:

<div id="ctl00_ContentPlaceHolder1_pnlPages" style="width: auto;"> <a id="ctl00_ContentPlaceHolder1_page1" class="aspNetDisabled" style="display:inline-block;font-weight:bold;height:22px;width:22px;background-image:url(../Images/count.png);">1</a><a id="ctl00_ContentPlaceHolder1_page2" href="javascript:__doPostBack('ctl00$ContentPlaceHolder1$page2','')" style="display:inline-block;height:22px;width:22px;background-image:url(../Images/count.png);">2</a><a id="ctl00_ContentPlaceHolder1_page3" href="javascript:__doPostBack('ctl00$ContentPlaceHolder1$page3','')" style="display:inline-block;height:22px;width:22px;background-image:url(../Images/count.png);">3</a><a id="ctl00_ContentPlaceHolder1_page4" href="javascript:__doPostBack('ctl00$ContentPlaceHolder1$page4','')" style="display:inline-block;height:22px;width:22px;background-image:url(../Images/count.png);">4</a><a id="ctl00_ContentPlaceHolder1_page5" href="javascript:__doPostBack('ctl00$ContentPlaceHolder1$page5','')" style="display:inline-block;height:22px;width:22px;background-image:url(../Images/count.png);">5</a> </div>

First off, we want to skip the disabled first page link (it has the aspNetDisabled class and no valid href). We can use either CSS or XPath selectors to grab only the active links with an href attribute:

  • CSS Selector: #ctl00_ContentPlaceHolder1_pnlPages a[href]
  • XPath: //div[@id='ctl00_ContentPlaceHolder1_pnlPages']/a[@href]
2. Extract the href Attribute Values

Once we've selected the links, extracting their href values is straightforward in Scrapy. Here's how to do this in your spider's parse method:

def parse(self, response):
    # Use CSS selector to get active pagination links
    pagination_links = response.css("#ctl00_ContentPlaceHolder1_pnlPages a[href]")
    
    # Or use XPath if that's your preference
    # pagination_links = response.xpath("//div[@id='ctl00_ContentPlaceHolder1_pnlPages']/a[@href]")
    
    # Extract all href attributes into a list
    href_list = pagination_links.xpath("@href").getall()
    
    # Verify the list (you can log this instead of print for production)
    print("Extracted JavaScript hrefs:", href_list)
3. Parse the __doPostBack Parameters (Critical for Loading Pages)

These href values aren't direct URLs—they're calls to ASP.NET's __doPostBack function, which triggers a POST request to load the next page. We need to extract the two parameters from this function call to build a valid request.

We can use a regex to pull out these parameters, then construct a FormRequest (Scrapy's built-in way to handle POST forms):

import re
import scrapy

def parse(self, response):
    pagination_links = response.css("#ctl00_ContentPlaceHolder1_pnlPages a[href]")
    
    # Loop through each active link
    for link in pagination_links:
        href = link.xpath("@href").get()
        # Regex to capture the two arguments inside __doPostBack
        match = re.search(r"__doPostBack\('(.*?)','(.*?)'\)", href)
        
        if match:
            event_target = match.group(1)
            event_argument = match.group(2)
            page_number = link.xpath("text()").get()
            
            # ASP.NET pages usually require hidden form fields like VIEWSTATE
            viewstate = response.css("#__VIEWSTATE::attr(value)").get()
            viewstate_generator = response.css("#__VIEWSTATEGENERATOR::attr(value)").get()
            
            # Yield a FormRequest to load the next page
            yield scrapy.FormRequest(
                url=response.url,
                formdata={
                    "__EVENTTARGET": event_target,
                    "__EVENTARGUMENT": event_argument,
                    "__VIEWSTATE": viewstate,
                    "__VIEWSTATEGENERATOR": viewstate_generator
                    # Add any other required form fields you find in the page's HTML
                },
                callback=self.parse_paginated_content,
                meta={"page_number": page_number}  # Pass page number to callback
            )

def parse_paginated_content(self, response):
    # This is where you extract data from each paginated page
    current_page = response.meta["page_number"]
    print(f"Parsing content from page {current_page}")
    
    # Example: Extract all titles (replace with your actual data extraction logic)
    # titles = response.css(".product-title::text").getall()
    # for title in titles:
    #     yield {"page": current_page, "product_title": title.strip()}
4. Key Notes for Success
  • Hidden Form Fields: ASP.NET relies on fields like __VIEWSTATE and __VIEWSTATEGENERATOR to validate requests. Always extract these from the current page's HTML—they change with each request.
  • Loop Handling: The for loop in the parse method automatically iterates over every active pagination link, sending a request for each page.
  • Callback Function: Using parse_paginated_content keeps your code organized—separating link extraction from content parsing.

内容的提问来源于stack exchange,提问作者rajeshbojja

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:06:37