如何使用requests模块获取目标网页分页中的剩余名称数据(替代Selenium方案)
Great question! The issue you're facing is super common with ASP.NET WebForms sites—they rely on hidden state fields like __VIEWSTATE, __VIEWSTATEGENERATOR, and __EVENTVALIDATION to handle postbacks (like pagination or "show all" actions). Your initial request works for the first page, but you need to capture and update these state fields for subsequent requests, plus include the right command parameters to trigger the action you want.
Let’s break down two solid solutions using requests—starting with the simpler "show all" approach (since it gets all 410 results in one go), then pagination if you prefer fetching page by page.
This mirrors what your Selenium code does, but entirely with requests. The "show all" link triggers a RadGrid command, so we just need to add the correct postback parameters to our payload.
Here’s the updated code:
import re import requests from bs4 import BeautifulSoup URL = "https://cci-online.org/CCI/Verify/CCI/Credential_Verification.aspx" params = { 'errorpath': '/CCI/Verify/CCI/Credential_Verification.aspx' } with requests.Session() as s: s.headers['User-Agent'] = 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/85.0.4183.102 Safari/537.36' # Step 1: Grab initial page data to get required keys and form fields r = s.get(URL) params['WebsiteKey'] = re.search(r"gWebsiteKey[^\\']+\\'(.*?)\\'", r.text).group(1) params['hkey'] = re.search(r"gHKey[^\\']+\\'(.*?)\\'", r.text).group(1) soup = BeautifulSoup(r.text, "lxml") payload = {i['name']: i.get('value', '') for i in soup.select('input[name]')} # Step 2: Submit the Alabama + United States filter payload['ctl01$TemplateBody$WebPartManager1$gwpciPeopleSearch$ciPeopleSearch$ResultsGrid$Sheet0$Input4$DropDown1'] = 'AL' payload['ctl01$TemplateBody$WebPartManager1$gwpciPeopleSearch$ciPeopleSearch$ResultsGrid$Sheet0$Input5$DropDown1'] = 'United States' r = s.post(URL, params=params, data=payload) soup = BeautifulSoup(r.text, "lxml") # Step 3: Update payload with fresh state fields from the first results page payload.update({ '__VIEWSTATE': soup.select_one('#__VIEWSTATE')['value'], '__VIEWSTATEGENERATOR': soup.select_one('#__VIEWSTATEGENERATOR')['value'], '__EVENTVALIDATION': soup.select_one('#__EVENTVALIDATION')['value'], }) # Step 4: Add parameters to trigger the "show all" command # Dynamically extract the target control ID instead of hardcoding (more robust) show_all_target = re.search(r"__doPostBack\('(.*?)','ShowAll'\)", r.text).group(1) payload['__EVENTTARGET'] = show_all_target payload['__EVENTARGUMENT'] = 'ShowAll' # Step 5: Submit the request to load all results r = s.post(URL, params=params, data=payload) soup = BeautifulSoup(r.text, "lxml") # Extract and print all names for item in soup.select("table.rgMasterTable > tbody > tr a[title]"): print(item.text.strip())
Key Notes:
- The
__EVENTTARGETis pulled directly from the "show all" link's__doPostBackJavaScript call in the page source—this avoids hardcoding IDs that might change if the site updates. - Always reuse the same
Sessionobject to maintain cookies and session state between requests.
If you need to fetch results page by page instead, you can mimic clicking the "Next" button. Each page request requires updating the state fields and sending the correct page number via __EVENTARGUMENT.
Here’s how to do it:
import re import requests from bs4 import BeautifulSoup URL = "https://cci-online.org/CCI/Verify/CCI/Credential_Verification.aspx" params = { 'errorpath': '/CCI/Verify/CCI/Credential_Verification.aspx' } with requests.Session() as s: s.headers['User-Agent'] = 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/85.0.4183.102 Safari/537.36' # Initial setup r = s.get(URL) params['WebsiteKey'] = re.search(r"gWebsiteKey[^\\']+\\'(.*?)\\'", r.text).group(1) params['hkey'] = re.search(r"gHKey[^\\']+\\'(.*?)\\'", r.text).group(1) soup = BeautifulSoup(r.text, "lxml") payload = {i['name']: i.get('value', '') for i in soup.select('input[name]')} # Apply filters payload['ctl01$TemplateBody$WebPartManager1$gwpciPeopleSearch$ciPeopleSearch$ResultsGrid$Sheet0$Input4$DropDown1'] = 'AL' payload['ctl01$TemplateBody$WebPartManager1$gwpciPeopleSearch$ciPeopleSearch$ResultsGrid$Sheet0$Input5$DropDown1'] = 'United States' r = s.post(URL, params=params, data=payload) soup = BeautifulSoup(r.text, "lxml") # Extract first page names for item in soup.select("table.rgMasterTable > tbody > tr a[title]"): print(item.text.strip()) # Get total number of pages to know when to stop total_pages = int(soup.select_one(".rgNumPart a:last-child").text) current_page = 1 # Dynamically extract the pagination control target ID next_button_target = re.search(r"__doPostBack\('(.*?)','Page\\$2'\)", r.text).group(1) while current_page < total_pages: current_page += 1 # Update state fields with fresh values from the current page payload.update({ '__VIEWSTATE': soup.select_one('#__VIEWSTATE')['value'], '__VIEWSTATEGENERATOR': soup.select_one('#__VIEWSTATEGENERATOR')['value'], '__EVENTVALIDATION': soup.select_one('#__EVENTVALIDATION')['value'], }) # Set pagination command parameters payload['__EVENTTARGET'] = next_button_target payload['__EVENTARGUMENT'] = f'Page${current_page}' # Submit request for next page r = s.post(URL, params=params, data=payload) soup = BeautifulSoup(r.text, "lxml") # Extract names from the current page for item in soup.select("table.rgMasterTable > tbody > tr a[title]"): print(item.text.strip())
Key Notes:
- We extract the pagination control ID dynamically from the "Next" button's
__doPostBackcall to avoid broken code if the site changes control IDs. - The
__EVENTARGUMENTuses the formatPage$XwhereXis the page number you want to load.
内容的提问来源于stack exchange,提问作者MITHU

