如何在Python Scrapy中提取JS onclick属性的href列表并循环解析
Got it, let's walk through exactly how to extract those JavaScript-driven href values from your pagination HTML using Scrapy, and set up a loop to parse each page.
First, here's the pagination HTML we're working with:
<div id="ctl00_ContentPlaceHolder1_pnlPages" style="width: auto;"> <a id="ctl00_ContentPlaceHolder1_page1" class="aspNetDisabled" style="display:inline-block;font-weight:bold;height:22px;width:22px;background-image:url(../Images/count.png);">1</a><a id="ctl00_ContentPlaceHolder1_page2" href="javascript:__doPostBack('ctl00$ContentPlaceHolder1$page2','')" style="display:inline-block;height:22px;width:22px;background-image:url(../Images/count.png);">2</a><a id="ctl00_ContentPlaceHolder1_page3" href="javascript:__doPostBack('ctl00$ContentPlaceHolder1$page3','')" style="display:inline-block;height:22px;width:22px;background-image:url(../Images/count.png);">3</a><a id="ctl00_ContentPlaceHolder1_page4" href="javascript:__doPostBack('ctl00$ContentPlaceHolder1$page4','')" style="display:inline-block;height:22px;width:22px;background-image:url(../Images/count.png);">4</a><a id="ctl00_ContentPlaceHolder1_page5" href="javascript:__doPostBack('ctl00$ContentPlaceHolder1$page5','')" style="display:inline-block;height:22px;width:22px;background-image:url(../Images/count.png);">5</a> </div>
First off, we want to skip the disabled first page link (it has the aspNetDisabled class and no valid href). We can use either CSS or XPath selectors to grab only the active links with an href attribute:
- CSS Selector:
#ctl00_ContentPlaceHolder1_pnlPages a[href] - XPath:
//div[@id='ctl00_ContentPlaceHolder1_pnlPages']/a[@href]
href Attribute Values Once we've selected the links, extracting their href values is straightforward in Scrapy. Here's how to do this in your spider's parse method:
def parse(self, response): # Use CSS selector to get active pagination links pagination_links = response.css("#ctl00_ContentPlaceHolder1_pnlPages a[href]") # Or use XPath if that's your preference # pagination_links = response.xpath("//div[@id='ctl00_ContentPlaceHolder1_pnlPages']/a[@href]") # Extract all href attributes into a list href_list = pagination_links.xpath("@href").getall() # Verify the list (you can log this instead of print for production) print("Extracted JavaScript hrefs:", href_list)
__doPostBack Parameters (Critical for Loading Pages) These href values aren't direct URLs—they're calls to ASP.NET's __doPostBack function, which triggers a POST request to load the next page. We need to extract the two parameters from this function call to build a valid request.
We can use a regex to pull out these parameters, then construct a FormRequest (Scrapy's built-in way to handle POST forms):
import re import scrapy def parse(self, response): pagination_links = response.css("#ctl00_ContentPlaceHolder1_pnlPages a[href]") # Loop through each active link for link in pagination_links: href = link.xpath("@href").get() # Regex to capture the two arguments inside __doPostBack match = re.search(r"__doPostBack\('(.*?)','(.*?)'\)", href) if match: event_target = match.group(1) event_argument = match.group(2) page_number = link.xpath("text()").get() # ASP.NET pages usually require hidden form fields like VIEWSTATE viewstate = response.css("#__VIEWSTATE::attr(value)").get() viewstate_generator = response.css("#__VIEWSTATEGENERATOR::attr(value)").get() # Yield a FormRequest to load the next page yield scrapy.FormRequest( url=response.url, formdata={ "__EVENTTARGET": event_target, "__EVENTARGUMENT": event_argument, "__VIEWSTATE": viewstate, "__VIEWSTATEGENERATOR": viewstate_generator # Add any other required form fields you find in the page's HTML }, callback=self.parse_paginated_content, meta={"page_number": page_number} # Pass page number to callback ) def parse_paginated_content(self, response): # This is where you extract data from each paginated page current_page = response.meta["page_number"] print(f"Parsing content from page {current_page}") # Example: Extract all titles (replace with your actual data extraction logic) # titles = response.css(".product-title::text").getall() # for title in titles: # yield {"page": current_page, "product_title": title.strip()}
- Hidden Form Fields: ASP.NET relies on fields like
__VIEWSTATEand__VIEWSTATEGENERATORto validate requests. Always extract these from the current page's HTML—they change with each request. - Loop Handling: The
forloop in theparsemethod automatically iterates over every active pagination link, sending a request for each page. - Callback Function: Using
parse_paginated_contentkeeps your code organized—separating link extraction from content parsing.
内容的提问来源于stack exchange,提问作者rajeshbojja

