使用网络请求爬取网站时访问端点遭遇400错误(Bad Request)求助
Hey there, let's dig into why you're hitting that persistent 400 Bad Request error when accessing the https://www.2embed.ru/ajax/embed/play endpoint. Here are the most likely culprits and actionable fixes:
1. Stale/Invalid CSRF Token (_token parameter)
The _token in your params is a CSRF protection token—these are tightly tied to your active session and typically expire quickly or are single-use. The token you're using is almost certainly stale, so the server rejects the request as a potential cross-site attack.
Fix:
- First, use a
requests.Session()to visit the referrer page (https://www.2embed.ru/embed/tmdb/movie?id=299534). This maintains your session state. - Extract a fresh
_tokenfrom the page's HTML (look for a hidden<input>withname="_token"or a meta tag containing the token). Tools like BeautifulSoup make parsing this straightforward.
2. Mismatched/Expired Cookies
You're manually setting cookies that are likely expired or not synced with your token. Servers often link CSRF tokens to specific session cookies—using old cookies with a new token (or vice versa) will trigger a 400 error.
Fix:
- Ditch manual cookie definitions and use
requests.Session()instead. The session automatically persists cookies between requests, ensuring your cookies and token are aligned.
3. Malformed Header Values
Your headers use " (HTML entity quotes) instead of actual double quotes. When requests sends these, the server may misinterpret the header format, leading to a bad request.
Fix:
- Replace all
"entries with regular"characters. For example:'sec-ch-ua': '" Not A;Brand";v="99", "Chromium";v="99", "Google Chrome";v="99"',
4. Potential Resource ID Mismatch
Your referrer page targets movie ID 299534, but your request uses id=1615802 in params. Double-check that this ID corresponds to a valid stream/resource linked to the movie page you're scraping.
Example Modified Code
Here's how to adjust your code to address all these issues:
import requests from bs4 import BeautifulSoup # Use a session to maintain cookies and session state session = requests.Session() # Step 1: Visit the referrer page to get fresh cookies and CSRF token referer_url = "https://www.2embed.ru/embed/tmdb/movie?id=299534" referer_response = session.get(referer_url) soup = BeautifulSoup(referer_response.text, "html.parser") # Extract the CSRF token (adjust selector if the page structure changes) csrf_token = soup.find("input", {"name": "_token"})["value"] # Step 2: Define corrected headers headers = { 'authority': 'www.2embed.ru', 'sec-ch-ua': '" Not A;Brand";v="99", "Chromium";v="99", "Google Chrome";v="99"', 'accept': '*/*', 'x-requested-with': 'XMLHttpRequest', 'sec-ch-ua-mobile': '?1', 'user-agent': 'Mozilla/5.0 (Linux; Android 6.0; Nexus 5 Build/MRA58N) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.74 Mobile Safari/537.36', 'sec-ch-ua-platform': '"Android"', 'sec-fetch-site': 'same-origin', 'sec-fetch-mode': 'cors', 'sec-fetch-dest': 'empty', 'referer': referer_url, 'accept-language': 'en-GB,en;q=0.9,en-US;q=0.8', } # Step 3: Use fresh token and verified resource ID params = { 'id': '1615802', # Confirm this ID matches the stream on the referer page '_token': csrf_token, } # Step 4: Make the request with the session response = session.get('https://www.2embed.ru/ajax/embed/play', headers=headers, params=params) # Check results print(response.status_code) print(response.text)
Quick Extra Tips
- Mimic real user behavior: Add small delays between requests (use
time.sleep()) and rotate user-agents occasionally to avoid being flagged as a scraper. - If the site still blocks you, check for additional anti-scraping measures like IP bans or JavaScript-rendered content (you might need tools like Selenium for that).
内容的提问来源于stack exchange,提问作者Aryan Agarwal

