如何仅下载网页部分HTML?Python Requests中Range请求头失效问题
Hey there! Let's break down this problem. The reason your Range header isn't limiting the downloaded content is that not all web servers support partial content requests—especially dynamic content providers like Google.
Google's search pages are generated on-the-fly when you send a request; they aren't stored as static files on the server that can be split into byte ranges. So when you send a Range header, Google's server simply ignores it and sends the full page anyway.
Alternative 1: Stream the Response & Read Only What You Need
Instead of relying on the server to honor Range, use requests' streaming feature to stop downloading as soon as you have enough content. Here's how to adjust your code:
import requests query = 'movie' size = 10 start = 0 USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" session = requests.Session() google_url = 'https://www.google.com/search?q={}&num={}&start={}'.format(query, size, start) # Enable streaming with stream=True response = session.get(google_url, verify=False, headers={'User-Agent': USER_AGENT}, stream=True) # Define how much content you want to grab (e.g., first 2000 bytes) max_bytes = 2000 partial_content = b'' # Iterate over response chunks until we hit our limit for chunk in response.iter_content(chunk_size=512): if chunk: # Filter out keep-alive empty chunks partial_content += chunk if len(partial_content) >= max_bytes: break # Convert bytes to text (handle encoding properly) partial_html = partial_content.decode(response.encoding or 'utf-8', errors='ignore') print(partial_html)
This works because streaming lets you receive the response in small chunks. You can even tweak it to stop reading once you hit a specific HTML element (like the opening tag of search results) by checking the chunk content, rather than just limiting by byte count.
Alternative 2: Target Specific Content Directly (If Possible)
If your goal is to extract specific data (like Google search results) instead of grabbing random partial HTML, consider:
- Scanning the early parts of the page for embedded scripts or JSON-LD blocks that contain the data you need. You can stop downloading as soon as you locate this section.
- Using Google's Custom Search API (note: this requires an API key and has usage limits) if you prefer a structured approach over raw HTML scraping.
Important Notes
- Respect Website Policies: Always check
robots.txtand avoid aggressive scraping that could get your IP blocked. Add reasonable delays between requests and use a realistic User-Agent string. - Encoding Handling: Make sure to use the page's correct encoding (you can retrieve it from
response.encoding). Usingerrors='ignore'helps avoid crashes from incomplete byte sequences. - Dynamic Content: If the page relies on JavaScript to load content,
requestswon't capture it even with streaming. In that case, tools like Selenium would be needed—but partial downloads aren't feasible there, since you need the full page to render JS.
内容的提问来源于stack exchange,提问作者hamid

