如何用Python实现网页下载:支持离线查看与HTML代码解析
Great question! Replicating the browser's "Save Page As" functionality in Python—so you can view the page offline and sift through its HTML for specific info—requires handling two core tasks: grabbing the full page (including all linked assets like CSS, images, and JavaScript) and parsing the HTML to extract what you need. Let’s break this down with practical code examples:
When you use "Save Page As" in a browser, it saves the HTML file plus a folder of linked resources, and updates the HTML to point to those local files. Here’s how to mimic that:
First, install the required libraries if you haven’t already:
pip install requests beautifulsoup4
Then use this function to download and save the page:
import os import requests from urllib.parse import urljoin, urlparse from bs4 import BeautifulSoup def save_page_offline(url, save_dir="saved_page"): # Create a folder to store the page and assets os.makedirs(save_dir, exist_ok=True) # Mimic a browser request to avoid being blocked headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) response.raise_for_status() # Throw an error if the request fails soup = BeautifulSoup(response.text, "html.parser") # Find all tags that link to external resources resource_tags = soup.find_all(["link", "script", "img", "source", "iframe"]) for tag in resource_tags: # Get the attribute that holds the resource URL (href for links, src for others) attr = "href" if tag.name in ["link", "a"] else "src" resource_url = tag.get(attr) if not resource_url: continue # Skip if the tag has no resource URL # Convert relative URLs (like "/styles.css") to absolute URLs absolute_url = urljoin(url, resource_url) # Optional: Only download resources from the same domain as the page # Remove this check if you want to download external assets too if urlparse(absolute_url).netloc != urlparse(url).netloc: continue # Get a local filename for the resource filename = os.path.join(save_dir, os.path.basename(urlparse(absolute_url).path)) # Download the resource and save it locally try: resource_response = requests.get(absolute_url, headers=headers) resource_response.raise_for_status() with open(filename, "wb") as f: f.write(resource_response.content) # Update the HTML tag to point to the local file instead of the web URL tag[attr] = os.path.basename(urlparse(absolute_url).path) except Exception as e: print(f"Couldn't download {absolute_url}: {str(e)}") # Save the modified HTML file html_file = os.path.join(save_dir, "index.html") with open(html_file, "w", encoding="utf-8") as f: f.write(str(soup)) print(f"Page saved offline to {html_file}") return html_file
Now that you have the saved HTML (or even the raw response from the web), you can use BeautifulSoup to parse and filter out the info you need. Here’s an example function:
def extract_html_info(html_path): with open(html_path, "r", encoding="utf-8") as f: soup = BeautifulSoup(f.read(), "html.parser") # Example 1: Grab all heading tags (h1 to h6) print("=== All Headings ===") for heading in soup.find_all(["h1", "h2", "h3", "h4", "h5", "h6"]): print(f"- {heading.get_text(strip=True)}") # Example 2: Extract all links with a specific class # Replace "article-link" with your target class print("\n=== Target Links ===") for link in soup.find_all("a", class_="article-link"): link_text = link.get_text(strip=True) link_url = link.get("href") print(f"- Text: {link_text}, URL: {link_url}") # Example 3: Pull text from a specific section by ID # Replace "main-content" with your target ID print("\n=== Main Content ===") main_section = soup.find("div", id="main-content") if main_section: print(main_section.get_text(strip=True))
Use the two functions like this:
# Replace this with your target webpage URL target_url = "https://example.com" # Save the page offline saved_html = save_page_offline(target_url) # Extract info from the saved HTML extract_html_info(saved_html)
If the page uses JavaScript to load content (like many modern sites), requests alone won’t capture that. For dynamic pages, use requests-html (which renders JS) or Selenium:
Install requests-html:
pip install requests-html
Then use this function for dynamic pages:
from requests_html import HTMLSession def save_dynamic_page_offline(url, save_dir="saved_dynamic_page"): os.makedirs(save_dir, exist_ok=True) session = HTMLSession() response = session.get(url) response.html.render() # This runs the JS to load dynamic content # Save the fully rendered HTML html_file = os.path.join(save_dir, "index.html") with open(html_file, "w", encoding="utf-8") as f: f.write(response.html.html) # Optional: Add the resource download logic from earlier to save assets print(f"Dynamic page saved to {html_file}") return html_file
内容的提问来源于stack exchange,提问作者RageAgainstheMachine

