You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取遇403 Forbidden与Internal Server Error,求排查代码问题

Troubleshooting 403 Forbidden and 500 Internal Server Errors in Web Scraping Code

Let's break down the issues in your code and fix them step by step:

1. Why the first code block returns 403 Forbidden

Your HTTPConnection approach has critical mistakes in request formatting:

  • You’re encoding request headers into url_params and passing that as the POST body—this is backwards. The POST body should be your form_data, not the headers.
  • You didn’t set the required Content-Type: application/x-www-form-urlencoded header, which tells the server you’re sending form-encoded data.
  • The __VIEWSTATE and __EVENTVALIDATION values from your XPath query are lists (XPath returns all matching nodes), but you need to pass the actual string value (the first element of the list).
  • You didn’t preserve cookies from the initial GET request—many ASP.NET sites require cookies to maintain valid session state.

2. Why the second code block returns 500 Internal Server Error

Your requests.post() call has incorrect parameter handling:

  • You’re passing url_params as an extra positional argument, but requests.post() expects form data in the data parameter and headers in the headers parameter.
  • Again, __VIEWSTATE and __EVENTVALIDATION are being passed as lists instead of strings, which sends invalid form data and triggers a server error.
  • You’re not persisting cookies between the initial GET and POST request, which breaks session validation.

Fixed Code Implementation

Here's a corrected version using requests (it simplifies cookie handling and request formatting):

import requests
from bs4 import BeautifulSoup
from lxml.html import fromstring

base_url = "https://apct.gov.in/apportal/Search/ViewAPVATDealers.aspx"
header = {
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8",
    "Accept-Encoding": "gzip, deflate, br",
    "Accept-Language": "en-US,en;q=0.9",
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.132 Safari/537.36",
}

# Use Session to persist cookies between requests (critical for ASP.NET sites)
session = requests.Session()
session.headers.update(header)

# First GET request to fetch viewstate, event validation, and set session cookies
get_response = session.get(base_url)
get_response.raise_for_status()  # Throw error if GET request fails

# Parse page to extract required hidden form fields
page_tree = fromstring(get_response.content)
form_data = {
    "__EVENTTARGET": "",
    "__EVENTARGUMENT": "",
    "__LASTFOCUS": "",
    # Extract string values from XPath results (take first element of list)
    "__VIEWSTATE": page_tree.xpath('//*[@id="__VIEWSTATE"]/@value')[0],
    "__EVENTVALIDATION": page_tree.xpath('//*[@id="__EVENTVALIDATION"]/@value')[0],
    "ctl00$ContentPlaceHolder1$dropact": "LT",
    "ctl00$ContentPlaceHolder1$Ddl_Divisions": "GUNTUR",
    "ctl00$ContentPlaceHolder1$Ddl_Circles": "All Circles",
    "ctl00$ContentPlaceHolder1$ddlbusines": "Agent"
}

# Send POST request with valid form data
post_response = session.post(base_url, data=form_data)
print(post_response.status_code, post_response.reason)

# Optional: Parse the response content
soup = BeautifulSoup(post_response.content, "html.parser")
# Add your data extraction logic here

Additional Tips

  • Always use a Session for ASP.NET sites—they rely on session cookies and viewstate to validate requests.
  • Double-check all form fields: some ASP.NET pages include extra hidden fields (like __VIEWSTATEGENERATOR) that you may need to extract and include in form_data. Inspect the page’s HTML source to confirm.
  • If you still get 403 errors, add a Referer header set to the base URL to mimic a real browser flow.
  • Respect the site’s robots.txt and terms of service—avoid sending excessive requests in a short time.

内容的提问来源于stack exchange,提问作者gochi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:39:38