如何发起HTTP请求从服务器获取最小量数据(Python场景)
Great question! When you're trying to pull only the necessary data from BookMyShow (https://in.bookmyshow.com) to speed up your requests—instead of downloading the full HTML every time—there are several practical approaches tailored to this goal. Here are the most effective ones, with Python examples:
1. Target Mobile or AMP Endpoints
Since you mentioned mobile devices get smaller payloads, BookMyShow likely maintains a dedicated mobile version (usually at https://m.in.bookmyshow.com) or AMP (Accelerated Mobile Pages) endpoints. These are stripped down to essential content, so the response size is way smaller than the desktop site.
Use a real mobile user-agent to ensure the server serves the mobile variant:
import requests # Mimic an iPhone user-agent headers = { "User-Agent": "Mozilla/5.0 (iPhone; CPU iPhone OS 16_0 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.0 Mobile/15E148 Safari/604.1" } # Fetch the mobile site response = requests.get("https://m.in.bookmyshow.com", headers=headers) minimal_html = response.text # Parse to extract only what you need (example with BeautifulSoup) from bs4 import BeautifulSoup soup = BeautifulSoup(minimal_html, "html.parser") upcoming_movies = soup.find_all("div", class_="movie-card")
2. Tap into Unofficial Public APIs
Most modern sites (including BookMyShow) use backend APIs to load content dynamically. These APIs return JSON data, which is far lighter than HTML. To find them:
- Open your browser's DevTools (F12), switch to the Network tab
- Set the device to mobile mode, refresh the page
- Look for XHR/Fetch requests (usually to endpoints with
/api/in the URL)
Once you find a relevant API endpoint, call it directly to get only the data you need:
import requests headers = { "User-Agent": "Mozilla/5.0 (iPhone; CPU iPhone OS 16_0 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.0 Mobile/15E148 Safari/604.1", "Accept": "application/json" } # Example endpoint (replace with one you find via DevTools) response = requests.get("https://in.bookmyshow.com/api/movies/upcoming", headers=headers) movie_data = response.json() # Extract only the fields you care about trimmed_data = [{"title": movie["title"], "release_date": movie["releaseDate"]} for movie in movie_data["data"]]
Note: Unofficial APIs can change without warning, so handle errors gracefully. Also, avoid spamming requests to avoid getting blocked.
3. Use Range Headers (For Partial Content)
Some servers support the Range HTTP header, which lets you request only a portion of the response. This works if you only need content from the start/end of the page:
import requests headers = { "User-Agent": "Magic Browser", "Range": "bytes=0-2000" # Fetch first 2000 bytes } response = requests.get("https://in.bookmyshow.com", headers=headers) partial_content = response.text
Keep in mind: Not all servers support this, and you might end up with incomplete HTML if the content you need isn't in the range you requested.
4. Switch to requests Instead of urllib2
While not a data-reduction trick, using the requests library (instead of outdated urllib2) makes your code cleaner and adds useful features like connection pooling, timeouts, and session persistence—all of which can speed up your requests overall.
Final Recommendation
The best bet is to use the mobile endpoint first, since it's designed for minimal payloads. If you need structured data, hunting for an API endpoint will give you the smallest possible response size.
内容的提问来源于stack exchange,提问作者Nikhil Wagh

