使用Python修改下载的HTML文件:将资源相对路径替换为带域名的绝对路径
Hey there! Looking at your code, you've already nailed the resource fetching part, but you're missing the key step of replacing relative paths in the HTML with full absolute URLs and saving the updated file as index2.html. Let's fix that up with a complete, revised version of your code, plus breakdowns of the critical changes:
Revised Code
import os import requests from bs4 import BeautifulSoup as bs from urllib.parse import urljoin from pathlib import Path import urllib.request, urllib.error, urllib.parse from tldextract import extract # Fix directory setup (original code had directory set to None since mkdir returns no value) dir_path = r"\Documents\python\public_html" Path(dir_path).mkdir(parents=True, exist_ok=True) # Target webpage URL url = "https://es.sopranodesign.com/sei/login.do?customerId=270" # Initialize session with a browser-like user agent to avoid blocking session = requests.Session() session.headers["User-Agent"] = "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/44.0.2403.157 Safari/537.36" # Fetch and parse the HTML content response = session.get(url) soup = bs(response.content, "html.parser") # -------------------------- # Core Feature: Replace relative paths with absolute URLs # -------------------------- # Update script tags for script in soup.find_all("script", src=True): relative_src = script.attrs["src"] absolute_src = urljoin(url, relative_src) script["src"] = absolute_src # Update CSS link tags for link in soup.find_all("link", href=True): relative_href = link.attrs["href"] absolute_href = urljoin(url, relative_href) link["href"] = absolute_href # Update image tags for img in soup.find_all("img", src=True): relative_src = img.attrs["src"] absolute_src = urljoin(url, relative_src) img["src"] = absolute_src # Save the modified HTML as index2.html with open(os.path.join(dir_path, "index2.html"), "w", encoding="utf-8") as f: f.write(str(soup)) # -------------------------- # Fetch and save external resources (JS, CSS, Images) # -------------------------- # Collect absolute resource URLs from the modified HTML script_files = [script["src"] for script in soup.find_all("script", src=True)] css_files = [link["href"] for link in soup.find_all("link", href=True)] image_files = [img["src"] for img in soup.find_all("img", src=True)] # Print resource counts for verification print("Total script files in the page:", len(script_files)) print("Total CSS files in the page:", len(css_files)) print("Total images files in the page:", len(image_files)) # Save resource URLs to text files for reference with open(os.path.join(dir_path, "javascript_files.txt"), "w") as f: for js_file in script_files: print(js_file, file=f) with open(os.path.join(dir_path, "css_files.txt"), "w") as f: for css_file in css_files: print(css_file, file=f) with open(os.path.join(dir_path, "image_files.txt"), "w") as f: for image_file in image_files: print(image_file, file=f) # Download JS files with error handling for js_file in script_files: try: file_name = os.path.basename(js_file) file_path = os.path.join(dir_path, file_name) js_content = session.get(js_file).text with open(file_path, "w", encoding="utf-8") as f: f.write(js_content) except Exception as e: print(f"Failed to download {js_file}: {str(e)}") # Download CSS files with error handling for css_file in css_files: try: file_name = os.path.basename(css_file) file_path = os.path.join(dir_path, file_name) css_content = session.get(css_file).text with open(file_path, "w", encoding="utf-8") as f: f.write(css_content) except Exception as e: print(f"Failed to download {css_file}: {str(e)}") # Download images (use binary mode to avoid corruption) for image_file in image_files: try: file_name = os.path.basename(image_file) file_path = os.path.join(dir_path, file_name) image_content = session.get(image_file).content with open(file_path, "wb") as f: # wb = write binary f.write(image_content) except Exception as e: print(f"Failed to download {image_file}: {str(e)}")
Key Improvements & Explanations
- Direct Path Replacement: Instead of just collecting resource URLs, we modify the
src/hrefattributes directly in the BeautifulSoup object. Usingurljoin()ensures even relative paths like/static/...get converted to full domain URLs. - Fixed Directory Handling: Separated directory creation from variable assignment (your original code had
directoryset toNonesincemkdir()returns no value). - Proper Image Downloads: Images are binary files, so we use
session.get(image_file).contentand open files inwbmode—this prevents corrupt image files that would happen with text-mode writing. - Transparent Error Handling: Added try-except blocks with clear error messages for each resource download, so you can spot failed fetches instead of silently skipping them.
- Consistent Resource Collection: After updating the HTML, we pull absolute URLs directly from the modified soup object, ensuring full alignment between the HTML and downloaded resources.
内容的提问来源于stack exchange,提问作者tony michael
相关产品推荐
相关产品推荐

