You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现Python版wget:网页存为index.html,文件保留原名?

问题

我正在用Python重写Unix的wget命令做趣味项目,需求是:下载网页时保存为index.html,下载普通文件则保留原文件名,但当前代码无法实现这个逻辑。现有代码如下:

import os
import requests
user_input = "wget google.com"
parts = user_input.split()
if parts[0] == "wget":
    # Check if the user just typed in wget
    if user_input == "wget":
        print(
            "wget: missing URL\nUsage: wget [URL]..."
            )
    # If not, run wget
    else:
        try:
            # Store the url the user put in in a variable
            url = parts[1]
            # Get the file name from the url
            file_name = url.split("/")[-1]
            # Destination
            destination = os.path.join(os.getcwd(),file_name)
            # Checking if the user typed in a url with http or https. If they didn't, it will add http:// to the url.
            if not url.startswith("https://") and not url.startswith("http://"):
                url = "http://" + url
            # Send an HTTP GET request to the URL
            response = requests.get(url)
            # If the response status code is not 200, raise an exception
            response.raise_for_status()

            # Write the response content to the destination file
            with open(destination, 'wb') as f:
                f.write(response.content)
            print("'{}' saved".format(file_name))

        # Catching any exception
        except Exception as e:
            print("wget: An error occurred:", e)
            print("Response status code:", response.status_code)
            print("Response content:", response.content)

解决方案

要实现区分网页和普通文件的命名逻辑,可以从响应的Content-Type头和URL路径特征两方面判断:

  • 当URL以/结尾(通常对应网站根目录或目录页),或者响应的Content-Type包含text/html时,判定为网页,使用index.html作为保存文件名;
  • 其他情况保留原URL提取的文件名。

修改后的代码如下(关键修改处已标注):

import os
import requests
user_input = "wget google.com"
parts = user_input.split()
if parts[0] == "wget":
    if user_input == "wget":
        print("wget: missing URL\nUsage: wget [URL]...")
    else:
        try:
            url = parts[1]
            # 初始从URL提取文件名
            file_name = url.split("/")[-1]
            # 补全URL协议头
            if not url.startswith("https://") and not url.startswith("http://"):
                url = "http://" + url
            
            response = requests.get(url)
            response.raise_for_status()

            # --- 新增:判断是否为网页,调整文件名 ---
            is_webpage = False
            # 1. 检查URL是否以/结尾(目录页)
            if url.endswith("/"):
                is_webpage = True
            # 2. 检查响应Content-Type是否为HTML
            content_type = response.headers.get("Content-Type", "")
            if "text/html" in content_type.lower():
                is_webpage = True
            
            if is_webpage:
                file_name = "index.html"
            # --- 文件名调整结束 ---

            destination = os.path.join(os.getcwd(), file_name)
            with open(destination, 'wb') as f:
                f.write(response.content)
            print(f"'{file_name}' saved")

        except Exception as e:
            print("wget: An error occurred:", e)
            # 避免未定义response的情况
            if 'response' in locals():
                print("Response status code:", response.status_code)
                print("Response content:", response.content)

关键修改说明

  1. 在获取响应后,新增网页判断逻辑:结合URL结尾是否为/和响应的Content-Type头判定是否为网页;
  2. 若判定为网页,将文件名强制设为index.html;
  3. 异常处理中增加response变量存在性判断,避免报错时因未定义response导致二次错误。

内容的提问来源于stack exchange,提问作者pizzavitdit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 02:30:44