You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python修改下载的HTML文件:将资源相对路径替换为带域名的绝对路径

Hey there! Looking at your code, you've already nailed the resource fetching part, but you're missing the key step of replacing relative paths in the HTML with full absolute URLs and saving the updated file as index2.html. Let's fix that up with a complete, revised version of your code, plus breakdowns of the critical changes:

Revised Code
import os
import requests
from bs4 import BeautifulSoup as bs
from urllib.parse import urljoin
from pathlib import Path
import urllib.request, urllib.error, urllib.parse
from tldextract import extract

# Fix directory setup (original code had directory set to None since mkdir returns no value)
dir_path = r"\Documents\python\public_html"
Path(dir_path).mkdir(parents=True, exist_ok=True)

# Target webpage URL
url = "https://es.sopranodesign.com/sei/login.do?customerId=270"
# Initialize session with a browser-like user agent to avoid blocking
session = requests.Session()
session.headers["User-Agent"] = "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/44.0.2403.157 Safari/537.36"

# Fetch and parse the HTML content
response = session.get(url)
soup = bs(response.content, "html.parser")

# --------------------------
# Core Feature: Replace relative paths with absolute URLs
# --------------------------
# Update script tags
for script in soup.find_all("script", src=True):
    relative_src = script.attrs["src"]
    absolute_src = urljoin(url, relative_src)
    script["src"] = absolute_src

# Update CSS link tags
for link in soup.find_all("link", href=True):
    relative_href = link.attrs["href"]
    absolute_href = urljoin(url, relative_href)
    link["href"] = absolute_href

# Update image tags
for img in soup.find_all("img", src=True):
    relative_src = img.attrs["src"]
    absolute_src = urljoin(url, relative_src)
    img["src"] = absolute_src

# Save the modified HTML as index2.html
with open(os.path.join(dir_path, "index2.html"), "w", encoding="utf-8") as f:
    f.write(str(soup))

# --------------------------
# Fetch and save external resources (JS, CSS, Images)
# --------------------------
# Collect absolute resource URLs from the modified HTML
script_files = [script["src"] for script in soup.find_all("script", src=True)]
css_files = [link["href"] for link in soup.find_all("link", href=True)]
image_files = [img["src"] for img in soup.find_all("img", src=True)]

# Print resource counts for verification
print("Total script files in the page:", len(script_files))
print("Total CSS files in the page:", len(css_files))
print("Total images files in the page:", len(image_files))

# Save resource URLs to text files for reference
with open(os.path.join(dir_path, "javascript_files.txt"), "w") as f:
    for js_file in script_files:
        print(js_file, file=f)

with open(os.path.join(dir_path, "css_files.txt"), "w") as f:
    for css_file in css_files:
        print(css_file, file=f)

with open(os.path.join(dir_path, "image_files.txt"), "w") as f:
    for image_file in image_files:
        print(image_file, file=f)

# Download JS files with error handling
for js_file in script_files:
    try:
        file_name = os.path.basename(js_file)
        file_path = os.path.join(dir_path, file_name)
        js_content = session.get(js_file).text
        with open(file_path, "w", encoding="utf-8") as f:
            f.write(js_content)
    except Exception as e:
        print(f"Failed to download {js_file}: {str(e)}")

# Download CSS files with error handling
for css_file in css_files:
    try:
        file_name = os.path.basename(css_file)
        file_path = os.path.join(dir_path, file_name)
        css_content = session.get(css_file).text
        with open(file_path, "w", encoding="utf-8") as f:
            f.write(css_content)
    except Exception as e:
        print(f"Failed to download {css_file}: {str(e)}")

# Download images (use binary mode to avoid corruption)
for image_file in image_files:
    try:
        file_name = os.path.basename(image_file)
        file_path = os.path.join(dir_path, file_name)
        image_content = session.get(image_file).content
        with open(file_path, "wb") as f:  # wb = write binary
            f.write(image_content)
    except Exception as e:
        print(f"Failed to download {image_file}: {str(e)}")
Key Improvements & Explanations
  • Direct Path Replacement: Instead of just collecting resource URLs, we modify the src/href attributes directly in the BeautifulSoup object. Using urljoin() ensures even relative paths like /static/... get converted to full domain URLs.
  • Fixed Directory Handling: Separated directory creation from variable assignment (your original code had directory set to None since mkdir() returns no value).
  • Proper Image Downloads: Images are binary files, so we use session.get(image_file).content and open files in wb mode—this prevents corrupt image files that would happen with text-mode writing.
  • Transparent Error Handling: Added try-except blocks with clear error messages for each resource download, so you can spot failed fetches instead of silently skipping them.
  • Consistent Resource Collection: After updating the HTML, we pull absolute URLs directly from the modified soup object, ensuring full alignment between the HTML and downloaded resources.

内容的提问来源于stack exchange,提问作者tony michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:22:46