使用requests.session()爬取BookMyShow预订页:返回200但源码显示无效请求
Hey there! Let's break down why you're getting that "invalid request" message even with a 200 status code when scraping BookMyShow. Here are the most likely culprits to check:
BookMyShow, like most modern websites, checks request headers to verify if the request is coming from a legitimate browser. The default requests library sends a User-Agent like python-requests/x.x.x, which is a dead giveaway for a crawler.
You'll need to add realistic headers to mimic a browser. For example:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9', 'Referer': 'https://in.bookmyshow.com/' } response = requests.get(your_booking_url, headers=headers)
Pay attention to headers like User-Agent, Referer, and Accept-Language — these are often the first line of defense against scrapers.
Many booking pages rely on specific query parameters (like city ID, movie ID, or showtime slot ID) to return valid content. If you're hitting a generic URL without these parameters, the server won't know what content to serve and returns an "invalid request" response.
Use your browser's DevTools (Network tab) to inspect the actual request URL when you load the booking page manually. You'll likely see parameters like ?city=MUM&movieId=12345 appended to the URL — make sure to include these in your requests call.
BookMyShow uses cookies to track user sessions, even for unauthenticated users. If you don't carry over cookies from a initial visit to the homepage, the server might flag your request as suspicious.
Use requests.Session() to maintain a persistent session, which automatically handles cookies:
session = requests.Session() # First visit the homepage to get necessary cookies session.get('https://in.bookmyshow.com/', headers=headers) # Now make the request to the booking page using the same session response = session.get(your_booking_url, headers=headers)
Some parts of BookMyShow might require JavaScript execution to generate valid request tokens or session data. Since requests doesn't run JavaScript, you might be missing critical data that the server expects.
If headers and cookies don't fix the issue, try using a browser automation tool like Selenium or Playwright. These tools mimic real browser behavior, execute JavaScript, and can bypass many basic anti-scraping checks.
Even two requests might trigger rate limiting if they're sent too quickly. Try adding a small delay between requests using time.sleep(2) to mimic human browsing speed. If the issue persists, you might need to use proxy IPs to avoid being flagged as a crawler.
A quick tip: Compare your request's headers, parameters, and cookies against the browser's request (via DevTools) — any mismatch is a potential culprit.
内容的提问来源于stack exchange,提问作者Thedeadman619

