如何在Python Scrapy爬虫中发起多个表单请求?
Hey there! Since you're new to Python, Scrapy, and web scraping overall, let's walk through how to tweak your spider to send form requests for multiple names (like "Cruz") and scrape each of their results.
First, let's fix and expand the code you shared. The target page uses ASP.NET, which usually requires hidden form fields (like VIEWSTATE, VIEWSTATEGENERATOR) to process POST requests. We'll need to grab those first from the initial search page before sending our form submissions.
Step 1: Define Your Target Names List
Start by creating a list of all the names you want to search for—this makes it super easy to add or remove entries later.
Step 2: Rewrite start_requests to Fetch the Initial Page
We'll first request the search page to get the necessary form parameters, then loop through our names list to send a separate form request for each name.
Complete Working Example
import scrapy from scrapy.http import FormRequest class LoginSpider(scrapy.Spider): name = "CRSpider5" login_url = 'http://recordingsearch.car.elpasoco.com/rsui/opr/search.aspx' # List of names we want to scrape results for target_names = ["Cruz", "Smith", "Johnson"] def start_requests(self): # First, load the search page to grab hidden form fields yield scrapy.Request( url=self.login_url, callback=self.handle_search_page ) def handle_search_page(self, response): # Extract required hidden ASP.NET form fields viewstate = response.css('input#__VIEWSTATE::attr(value)').get() viewstate_generator = response.css('input#__VIEWSTATEGENERATOR::attr(value)').get() event_validation = response.css('input#__EVENTVALIDATION::attr(value)').get() # Loop through each name in our target list for name in self.target_names: # Build form data for the current name form_data = { '__VIEWSTATE': viewstate, '__VIEWSTATEGENERATOR': viewstate_generator, '__EVENTVALIDATION': event_validation, # Replace these with the actual input names from the page (use dev tools to check) 'txtLastName': name, 'btnSearch': 'Search' # The name/value of the search button } # Send a POST request for this name yield FormRequest( url=self.login_url, formdata=form_data, callback=self.parse_search_results, meta={'search_name': name} # Pass the name to the parser for reference ) def parse_search_results(self, response): # Retrieve the name we searched for from the meta data searched_name = response.meta['search_name'] # Parse results here—adjust selectors to match the actual page structure results = response.css('div.result-container .result-item') # Example selector for result in results: yield { 'searched_name': searched_name, 'document_title': result.css('h4.document-title::text').get(), 'filing_date': result.css('.filing-date::text').get(), # Add other fields you want to scrape }
Key Notes for Success:
- Hidden Form Fields: ASP.NET uses these to validate requests, so double-check their IDs with your browser's dev tools (right-click → Inspect) if they don't match the example.
- Form Data: Replace
txtLastNameandbtnSearchwith the actualnameattributes of the search input and button on the page—you'll find these via inspecting the form. - Meta Parameter: We use
metato pass the searched name to the parser, so we can link each scraped result to the name that generated it. - Testing: Start with one name first to confirm the setup works, then add more names to the list. Use
scrapy crawl CRSpider5 -o results.jsonto save your output to a JSON file for easy review.
内容的提问来源于stack exchange,提问作者Walt911

