如何用Python数据爬虫下载获权网站的全站HTML内容?
Hey Larry, since you’ve got full permission to scrape the target site, let’s break down how to build a crawler that grabs not just the main page but every linked page within the same domain—like https://www.dogs.com/about-us and all other related paths.
Step 1: Install Required Tools
First, grab the libraries we’ll need to fetch and parse pages:
pip install requests beautifulsoup4
Step 2: Build a Basic Recursive Crawler
We’ll track visited URLs to avoid looping back to the same pages and wasting requests. Here’s a working starting point:
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin, urlparse import time # Configure your target site BASE_URL = "https://www.dogs.com" visited_urls = set() def crawl_page(url): # Skip if we've already visited this URL or it's outside the target domain if url in visited_urls or urlparse(url).netloc != urlparse(BASE_URL).netloc: return print(f"Scraping: {url}") visited_urls.add(url) try: # Fetch the page (add a delay to be kind to the server) time.sleep(1) response = requests.get(url, timeout=10) response.raise_for_status() # Throw an error for 4xx/5xx status codes # Save the HTML content to a file # Create a safe filename from the URL (adjust this for complex paths) url_path = url.replace(BASE_URL, "").strip("/") filename = url_path.replace("/", "_") if url_path else "index" filename += ".html" with open(filename, "w", encoding="utf-8") as file: file.write(response.text) # Parse the page to find all internal links soup = BeautifulSoup(response.text, "html.parser") for link in soup.find_all("a", href=True): # Convert relative links to absolute URLs absolute_link = urljoin(BASE_URL, link["href"]) # Recursively crawl the linked page crawl_page(absolute_link) except requests.exceptions.RequestException as e: print(f"Failed to scrape {url}: {str(e)}") # Start crawling from the main site crawl_page(BASE_URL)
Step 3: Important Notes (Even With Authorization)
- Rate Limiting: The
time.sleep(1)adds a 1-second delay between requests to avoid overwhelming the server. Adjust this based on the site’s capacity. - Safe Filenames: The filename logic works for simple URLs, but for more complex paths (like those with query parameters), use a library like
python-slugifyto create valid filenames. - Dynamic Content: If the site loads content with JavaScript,
requestswon’t capture it. Use tools like Selenium or Playwright to render the full page. - Avoid Recursion Limits: For very large sites, recursive calls might hit Python’s recursion depth limit. Swap to an iterative approach using a queue instead:
from collections import deque visited_urls = set() queue = deque([BASE_URL]) while queue: url = queue.popleft() if url in visited_urls or urlparse(url).netloc != urlparse(BASE_URL).netloc: continue # Same fetch, save, parse logic as above... visited_urls.add(url) # Add new links to the queue instead of recursing soup = BeautifulSoup(response.text, "html.parser") for link in soup.find_all("a", href=True): absolute_link = urljoin(BASE_URL, link["href"]) if absolute_link not in visited_urls: queue.append(absolute_link)
- Respect Robots.txt: Even with authorization, it’s good practice to check the site’s
robots.txtfile to see if any paths are off-limits. You can use therobotexclusionrulesparserlibrary to automate this check.
内容的提问来源于stack exchange,提问作者Larry
相关产品推荐
相关产品推荐

