使用Python的request.post请求时遭遇403错误,添加Header仍未解决该如何处理?
Hey there, let's tackle this 403 Forbidden error together—adding headers is a good first step, but there's usually more to it when sites block automated requests. Here are some actionable steps to fix this:
Most sites don't just check the User-Agent—they validate a full set of headers to confirm you're a human visitor. Use your browser's DevTools (F12 > Network tab) to capture the exact headers from a real search request, then replicate them. For example:
headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 14_0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://target-site.com/search-page', # Match the page you'd navigate from before searching 'DNT': '1', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' }
Never make up headers—copy them exactly from your browser's real request.
Many sites require valid session cookies to process POST requests. Instead of manually handling cookies, use requests.Session() to automatically persist cookies from initial page loads:
import requests session = requests.Session() # First, visit the target site's homepage or search page to get session cookies session.get('https://target-site.com/', headers=headers) # Now send your POST request with the same session response = session.post( 'https://target-site.com/api/search', headers=headers, data={'query': 'your search term', 'other_param': 'value'} )
A 403 can also pop up if your request body doesn't match what the site expects. Check DevTools again:
- If the site sends data as
form-dataorx-www-form-urlencoded, use thedataparameter inpost(). - If it sends JSON, use the
jsonparameter instead ofdata:response = session.post(url, headers=headers, json={'query': 'your term'})
Double-check parameter names—even a typo like search_q instead of search_query can trigger a block.
Lots of sites embed a hidden CSRF token in search pages, which you must include in your POST payload. Fetch the token first by parsing the search page HTML:
from bs4 import BeautifulSoup # Get the search page search_page = session.get('https://target-site.com/search', headers=headers) soup = BeautifulSoup(search_page.text, 'html.parser') # Extract the token (adjust the selector to match the site's HTML) csrf_token = soup.find('input', {'name': 'csrf_token'})['value'] # Include the token in your payload payload = {'query': 'your term', 'csrf_token': csrf_token} response = session.post(url, headers=headers, data=payload)
Sites block rapid, repeated requests. Add delays between requests to mimic human behavior:
import time # After each request time.sleep(2) # Wait 2 seconds before the next request
If you're making many requests, consider rotating proxy IPs (use reputable paid proxies—free ones are often already blocked).
If the search results load dynamically via JavaScript, requests won't see them because it only fetches static HTML. Tools like Playwright or Selenium can simulate a real browser to load full content:
from playwright.sync_api import sync_playwright with sync_playwright() as p: browser = p.chromium.launch(headless=False) # Headless=True for background runs page = browser.new_page() page.goto('https://target-site.com/search') # Fill the search box and submit page.fill('input[name="query"]', 'your search term') page.click('button[type="submit"]') # Wait for results to load page.wait_for_load_state('networkidle') # Get the full rendered HTML results_html = page.content() browser.close()
内容的提问来源于stack exchange,提问作者Beherit

