You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python requests获取Chrome中「查看网页源代码」的内容(含YouTube示例)

Hey folks, let's walk through how to get that exact "View Page Source" content you see in Google Chrome using Python's requests library—then we'll dive into the YouTube-specific case since it has some quirks to watch out for.

How to Fetch "View Page Source" Content with Python's Requests Library

General Approach for Most Websites

First, let's clarify: when you click "View Page Source" in Chrome, you're looking at the raw HTML that the server sends directly to your browser—not the dynamically rendered DOM after JavaScript runs. That's key because requests can grab this raw content easily, as long as you mimic a real browser's request.

Here's a step-by-step breakdown:

  • Install requests first if you haven't: run pip install requests in your terminal.
  • Use realistic request headers: Most websites block requests that don't look like they're coming from a browser. At minimum, you'll need a User-Agent string to mimic Chrome.
  • Send the GET request and retrieve the response text.

Example code:

import requests

# Replace with your target URL
target_url = "https://example.com"

# Mimic a Chrome browser's request headers
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

try:
    # Send the GET request
    response = requests.get(target_url, headers=headers)
    # Check if the request was successful (status code 200)
    response.raise_for_status()
    
    # Print the source code directly, or save to a file
    print(response.text)
    # To save to a file:
    # with open("page_source.html", "w", encoding="utf-8") as f:
    #     f.write(response.text)
except requests.exceptions.RequestException as e:
    print(f"Error fetching page: {e}")

Specific Case: Fetching YouTube's Page Source

YouTube has stricter anti-bot measures, so you'll need to expand your headers a bit to avoid getting blocked or receiving a stripped-down version of the page.

The main additions are Accept-Language (to match your browser's language setting) and a more detailed Accept header. Here's a working example:

import requests

url = "https://youtube.com"

# More comprehensive headers to mimic a real Chrome session
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept-Language": "en-US,en;q=0.9",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8",
    "Upgrade-Insecure-Requests": "1"
}

try:
    response = requests.get(url, headers=headers)
    response.raise_for_status()
    
    # Save the source code to a file for easier inspection
    with open("youtube_source.html", "w", encoding="utf-8") as f:
        f.write(response.text)
    
    print("Successfully fetched YouTube's page source! Check the youtube_source.html file.")
except requests.exceptions.RequestException as e:
    print(f"Request failed: {e}")

Important Notes for YouTube:

  • Regional Differences: YouTube serves different content based on your location. If you're getting a region-locked version, you might need to add a Cookie header with your region preferences, but that's more complex. For basic source code, the headers above should work.
  • Dynamic Content: Remember, this only gets the initial page source (what you see in "View Page Source"). If you need content that loads after JavaScript runs (like video comments or dynamically generated sections), you'll need a tool like Selenium or Playwright to render the page fully.
  • Header Updates: Chrome's User-Agent string changes with updates. If your request starts failing, check your browser's actual User-Agent (search "what's my user agent" in Chrome) and update the string in your code.

内容的提问来源于stack exchange,提问作者dulatus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:13:05