You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何仅下载网页部分HTML?Python Requests中Range请求头失效问题

Why Your Range Header Isn't Working & Alternatives for Partial HTML Downloads

Hey there! Let's break down this problem. The reason your Range header isn't limiting the downloaded content is that not all web servers support partial content requests—especially dynamic content providers like Google.

Google's search pages are generated on-the-fly when you send a request; they aren't stored as static files on the server that can be split into byte ranges. So when you send a Range header, Google's server simply ignores it and sends the full page anyway.

Alternative 1: Stream the Response & Read Only What You Need

Instead of relying on the server to honor Range, use requests' streaming feature to stop downloading as soon as you have enough content. Here's how to adjust your code:

import requests

query = 'movie'
size = 10
start = 0
USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"

session = requests.Session()
google_url = 'https://www.google.com/search?q={}&num={}&start={}'.format(query, size, start)

# Enable streaming with stream=True
response = session.get(google_url, verify=False, headers={'User-Agent': USER_AGENT}, stream=True)

# Define how much content you want to grab (e.g., first 2000 bytes)
max_bytes = 2000
partial_content = b''

# Iterate over response chunks until we hit our limit
for chunk in response.iter_content(chunk_size=512):
    if chunk:  # Filter out keep-alive empty chunks
        partial_content += chunk
        if len(partial_content) >= max_bytes:
            break

# Convert bytes to text (handle encoding properly)
partial_html = partial_content.decode(response.encoding or 'utf-8', errors='ignore')
print(partial_html)

This works because streaming lets you receive the response in small chunks. You can even tweak it to stop reading once you hit a specific HTML element (like the opening tag of search results) by checking the chunk content, rather than just limiting by byte count.

Alternative 2: Target Specific Content Directly (If Possible)

If your goal is to extract specific data (like Google search results) instead of grabbing random partial HTML, consider:

  • Scanning the early parts of the page for embedded scripts or JSON-LD blocks that contain the data you need. You can stop downloading as soon as you locate this section.
  • Using Google's Custom Search API (note: this requires an API key and has usage limits) if you prefer a structured approach over raw HTML scraping.

Important Notes

  • Respect Website Policies: Always check robots.txt and avoid aggressive scraping that could get your IP blocked. Add reasonable delays between requests and use a realistic User-Agent string.
  • Encoding Handling: Make sure to use the page's correct encoding (you can retrieve it from response.encoding). Using errors='ignore' helps avoid crashes from incomplete byte sequences.
  • Dynamic Content: If the page relies on JavaScript to load content, requests won't capture it even with streaming. In that case, tools like Selenium would be needed—but partial downloads aren't feasible there, since you need the full page to render JS.

内容的提问来源于stack exchange,提问作者hamid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:48:51