网页爬取遇403 Forbidden与Internal Server Error,求排查代码问题
Troubleshooting 403 Forbidden and 500 Internal Server Errors in Web Scraping Code
Let's break down the issues in your code and fix them step by step:
1. Why the first code block returns 403 Forbidden
Your HTTPConnection approach has critical mistakes in request formatting:
- You’re encoding request headers into
url_paramsand passing that as the POST body—this is backwards. The POST body should be yourform_data, not the headers. - You didn’t set the required
Content-Type: application/x-www-form-urlencodedheader, which tells the server you’re sending form-encoded data. - The
__VIEWSTATEand__EVENTVALIDATIONvalues from your XPath query are lists (XPath returns all matching nodes), but you need to pass the actual string value (the first element of the list). - You didn’t preserve cookies from the initial GET request—many ASP.NET sites require cookies to maintain valid session state.
2. Why the second code block returns 500 Internal Server Error
Your requests.post() call has incorrect parameter handling:
- You’re passing
url_paramsas an extra positional argument, butrequests.post()expects form data in thedataparameter and headers in theheadersparameter. - Again,
__VIEWSTATEand__EVENTVALIDATIONare being passed as lists instead of strings, which sends invalid form data and triggers a server error. - You’re not persisting cookies between the initial GET and POST request, which breaks session validation.
Fixed Code Implementation
Here's a corrected version using requests (it simplifies cookie handling and request formatting):
import requests from bs4 import BeautifulSoup from lxml.html import fromstring base_url = "https://apct.gov.in/apportal/Search/ViewAPVATDealers.aspx" header = { "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8", "Accept-Encoding": "gzip, deflate, br", "Accept-Language": "en-US,en;q=0.9", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.132 Safari/537.36", } # Use Session to persist cookies between requests (critical for ASP.NET sites) session = requests.Session() session.headers.update(header) # First GET request to fetch viewstate, event validation, and set session cookies get_response = session.get(base_url) get_response.raise_for_status() # Throw error if GET request fails # Parse page to extract required hidden form fields page_tree = fromstring(get_response.content) form_data = { "__EVENTTARGET": "", "__EVENTARGUMENT": "", "__LASTFOCUS": "", # Extract string values from XPath results (take first element of list) "__VIEWSTATE": page_tree.xpath('//*[@id="__VIEWSTATE"]/@value')[0], "__EVENTVALIDATION": page_tree.xpath('//*[@id="__EVENTVALIDATION"]/@value')[0], "ctl00$ContentPlaceHolder1$dropact": "LT", "ctl00$ContentPlaceHolder1$Ddl_Divisions": "GUNTUR", "ctl00$ContentPlaceHolder1$Ddl_Circles": "All Circles", "ctl00$ContentPlaceHolder1$ddlbusines": "Agent" } # Send POST request with valid form data post_response = session.post(base_url, data=form_data) print(post_response.status_code, post_response.reason) # Optional: Parse the response content soup = BeautifulSoup(post_response.content, "html.parser") # Add your data extraction logic here
Additional Tips
- Always use a
Sessionfor ASP.NET sites—they rely on session cookies and viewstate to validate requests. - Double-check all form fields: some ASP.NET pages include extra hidden fields (like
__VIEWSTATEGENERATOR) that you may need to extract and include inform_data. Inspect the page’s HTML source to confirm. - If you still get 403 errors, add a
Refererheader set to the base URL to mimic a real browser flow. - Respect the site’s
robots.txtand terms of service—avoid sending excessive requests in a short time.
内容的提问来源于stack exchange,提问作者gochi
相关产品推荐
相关产品推荐

