如何使用Python requests获取Chrome中「查看网页源代码」的内容(含YouTube示例)
Hey folks, let's walk through how to get that exact "View Page Source" content you see in Google Chrome using Python's requests library—then we'll dive into the YouTube-specific case since it has some quirks to watch out for.
General Approach for Most Websites
First, let's clarify: when you click "View Page Source" in Chrome, you're looking at the raw HTML that the server sends directly to your browser—not the dynamically rendered DOM after JavaScript runs. That's key because requests can grab this raw content easily, as long as you mimic a real browser's request.
Here's a step-by-step breakdown:
- Install requests first if you haven't: run
pip install requestsin your terminal. - Use realistic request headers: Most websites block requests that don't look like they're coming from a browser. At minimum, you'll need a
User-Agentstring to mimic Chrome. - Send the GET request and retrieve the response text.
Example code:
import requests # Replace with your target URL target_url = "https://example.com" # Mimic a Chrome browser's request headers headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } try: # Send the GET request response = requests.get(target_url, headers=headers) # Check if the request was successful (status code 200) response.raise_for_status() # Print the source code directly, or save to a file print(response.text) # To save to a file: # with open("page_source.html", "w", encoding="utf-8") as f: # f.write(response.text) except requests.exceptions.RequestException as e: print(f"Error fetching page: {e}")
Specific Case: Fetching YouTube's Page Source
YouTube has stricter anti-bot measures, so you'll need to expand your headers a bit to avoid getting blocked or receiving a stripped-down version of the page.
The main additions are Accept-Language (to match your browser's language setting) and a more detailed Accept header. Here's a working example:
import requests url = "https://youtube.com" # More comprehensive headers to mimic a real Chrome session headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept-Language": "en-US,en;q=0.9", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8", "Upgrade-Insecure-Requests": "1" } try: response = requests.get(url, headers=headers) response.raise_for_status() # Save the source code to a file for easier inspection with open("youtube_source.html", "w", encoding="utf-8") as f: f.write(response.text) print("Successfully fetched YouTube's page source! Check the youtube_source.html file.") except requests.exceptions.RequestException as e: print(f"Request failed: {e}")
Important Notes for YouTube:
- Regional Differences: YouTube serves different content based on your location. If you're getting a region-locked version, you might need to add a
Cookieheader with your region preferences, but that's more complex. For basic source code, the headers above should work. - Dynamic Content: Remember, this only gets the initial page source (what you see in "View Page Source"). If you need content that loads after JavaScript runs (like video comments or dynamically generated sections), you'll need a tool like Selenium or Playwright to render the page fully.
- Header Updates: Chrome's
User-Agentstring changes with updates. If your request starts failing, check your browser's actualUser-Agent(search "what's my user agent" in Chrome) and update the string in your code.
内容的提问来源于stack exchange,提问作者dulatus

