You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python数据爬虫下载获权网站的全站HTML内容?

How to Recursively Scrape All HTML Content of a Website with Python (With Full Authorization)

Hey Larry, since you’ve got full permission to scrape the target site, let’s break down how to build a crawler that grabs not just the main page but every linked page within the same domain—like https://www.dogs.com/about-us and all other related paths.

Step 1: Install Required Tools

First, grab the libraries we’ll need to fetch and parse pages:

pip install requests beautifulsoup4

Step 2: Build a Basic Recursive Crawler

We’ll track visited URLs to avoid looping back to the same pages and wasting requests. Here’s a working starting point:

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse
import time

# Configure your target site
BASE_URL = "https://www.dogs.com"
visited_urls = set()

def crawl_page(url):
    # Skip if we've already visited this URL or it's outside the target domain
    if url in visited_urls or urlparse(url).netloc != urlparse(BASE_URL).netloc:
        return

    print(f"Scraping: {url}")
    visited_urls.add(url)

    try:
        # Fetch the page (add a delay to be kind to the server)
        time.sleep(1)
        response = requests.get(url, timeout=10)
        response.raise_for_status()  # Throw an error for 4xx/5xx status codes

        # Save the HTML content to a file
        # Create a safe filename from the URL (adjust this for complex paths)
        url_path = url.replace(BASE_URL, "").strip("/")
        filename = url_path.replace("/", "_") if url_path else "index"
        filename += ".html"
        
        with open(filename, "w", encoding="utf-8") as file:
            file.write(response.text)

        # Parse the page to find all internal links
        soup = BeautifulSoup(response.text, "html.parser")
        for link in soup.find_all("a", href=True):
            # Convert relative links to absolute URLs
            absolute_link = urljoin(BASE_URL, link["href"])
            # Recursively crawl the linked page
            crawl_page(absolute_link)

    except requests.exceptions.RequestException as e:
        print(f"Failed to scrape {url}: {str(e)}")

# Start crawling from the main site
crawl_page(BASE_URL)

Step 3: Important Notes (Even With Authorization)

  • Rate Limiting: The time.sleep(1) adds a 1-second delay between requests to avoid overwhelming the server. Adjust this based on the site’s capacity.
  • Safe Filenames: The filename logic works for simple URLs, but for more complex paths (like those with query parameters), use a library like python-slugify to create valid filenames.
  • Dynamic Content: If the site loads content with JavaScript, requests won’t capture it. Use tools like Selenium or Playwright to render the full page.
  • Avoid Recursion Limits: For very large sites, recursive calls might hit Python’s recursion depth limit. Swap to an iterative approach using a queue instead:
from collections import deque

visited_urls = set()
queue = deque([BASE_URL])

while queue:
    url = queue.popleft()
    if url in visited_urls or urlparse(url).netloc != urlparse(BASE_URL).netloc:
        continue
    
    # Same fetch, save, parse logic as above...
    visited_urls.add(url)
    
    # Add new links to the queue instead of recursing
    soup = BeautifulSoup(response.text, "html.parser")
    for link in soup.find_all("a", href=True):
        absolute_link = urljoin(BASE_URL, link["href"])
        if absolute_link not in visited_urls:
            queue.append(absolute_link)
  • Respect Robots.txt: Even with authorization, it’s good practice to check the site’s robots.txt file to see if any paths are off-limits. You can use the robotexclusionrulesparser library to automate this check.

内容的提问来源于stack exchange,提问作者Larry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:57:17