You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python实现网页下载:支持离线查看与HTML代码解析

Great question! Replicating the browser's "Save Page As" functionality in Python—so you can view the page offline and sift through its HTML for specific info—requires handling two core tasks: grabbing the full page (including all linked assets like CSS, images, and JavaScript) and parsing the HTML to extract what you need. Let’s break this down with practical code examples:

1. Download the Full Page for Offline Viewing

When you use "Save Page As" in a browser, it saves the HTML file plus a folder of linked resources, and updates the HTML to point to those local files. Here’s how to mimic that:

First, install the required libraries if you haven’t already:

pip install requests beautifulsoup4

Then use this function to download and save the page:

import os
import requests
from urllib.parse import urljoin, urlparse
from bs4 import BeautifulSoup

def save_page_offline(url, save_dir="saved_page"):
    # Create a folder to store the page and assets
    os.makedirs(save_dir, exist_ok=True)
    
    # Mimic a browser request to avoid being blocked
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    response = requests.get(url, headers=headers)
    response.raise_for_status()  # Throw an error if the request fails
    soup = BeautifulSoup(response.text, "html.parser")
    
    # Find all tags that link to external resources
    resource_tags = soup.find_all(["link", "script", "img", "source", "iframe"])
    
    for tag in resource_tags:
        # Get the attribute that holds the resource URL (href for links, src for others)
        attr = "href" if tag.name in ["link", "a"] else "src"
        resource_url = tag.get(attr)
        
        if not resource_url:
            continue  # Skip if the tag has no resource URL
        
        # Convert relative URLs (like "/styles.css") to absolute URLs
        absolute_url = urljoin(url, resource_url)
        
        # Optional: Only download resources from the same domain as the page
        # Remove this check if you want to download external assets too
        if urlparse(absolute_url).netloc != urlparse(url).netloc:
            continue
        
        # Get a local filename for the resource
        filename = os.path.join(save_dir, os.path.basename(urlparse(absolute_url).path))
        
        # Download the resource and save it locally
        try:
            resource_response = requests.get(absolute_url, headers=headers)
            resource_response.raise_for_status()
            with open(filename, "wb") as f:
                f.write(resource_response.content)
            
            # Update the HTML tag to point to the local file instead of the web URL
            tag[attr] = os.path.basename(urlparse(absolute_url).path)
        except Exception as e:
            print(f"Couldn't download {absolute_url}: {str(e)}")
    
    # Save the modified HTML file
    html_file = os.path.join(save_dir, "index.html")
    with open(html_file, "w", encoding="utf-8") as f:
        f.write(str(soup))
    
    print(f"Page saved offline to {html_file}")
    return html_file
2. Extract Information from the HTML

Now that you have the saved HTML (or even the raw response from the web), you can use BeautifulSoup to parse and filter out the info you need. Here’s an example function:

def extract_html_info(html_path):
    with open(html_path, "r", encoding="utf-8") as f:
        soup = BeautifulSoup(f.read(), "html.parser")
    
    # Example 1: Grab all heading tags (h1 to h6)
    print("=== All Headings ===")
    for heading in soup.find_all(["h1", "h2", "h3", "h4", "h5", "h6"]):
        print(f"- {heading.get_text(strip=True)}")
    
    # Example 2: Extract all links with a specific class
    # Replace "article-link" with your target class
    print("\n=== Target Links ===")
    for link in soup.find_all("a", class_="article-link"):
        link_text = link.get_text(strip=True)
        link_url = link.get("href")
        print(f"- Text: {link_text}, URL: {link_url}")
    
    # Example 3: Pull text from a specific section by ID
    # Replace "main-content" with your target ID
    print("\n=== Main Content ===")
    main_section = soup.find("div", id="main-content")
    if main_section:
        print(main_section.get_text(strip=True))
Putting It All Together

Use the two functions like this:

# Replace this with your target webpage URL
target_url = "https://example.com"

# Save the page offline
saved_html = save_page_offline(target_url)

# Extract info from the saved HTML
extract_html_info(saved_html)
Handling Dynamic Content

If the page uses JavaScript to load content (like many modern sites), requests alone won’t capture that. For dynamic pages, use requests-html (which renders JS) or Selenium:

Install requests-html:

pip install requests-html

Then use this function for dynamic pages:

from requests_html import HTMLSession

def save_dynamic_page_offline(url, save_dir="saved_dynamic_page"):
    os.makedirs(save_dir, exist_ok=True)
    session = HTMLSession()
    response = session.get(url)
    response.html.render()  # This runs the JS to load dynamic content
    
    # Save the fully rendered HTML
    html_file = os.path.join(save_dir, "index.html")
    with open(html_file, "w", encoding="utf-8") as f:
        f.write(response.html.html)
    
    # Optional: Add the resource download logic from earlier to save assets
    print(f"Dynamic page saved to {html_file}")
    return html_file

内容的提问来源于stack exchange,提问作者RageAgainstheMachine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:17:32