如何实现Python版wget:网页存为index.html,文件保留原名?
问题
我正在用Python重写Unix的wget命令做趣味项目,需求是:下载网页时保存为index.html,下载普通文件则保留原文件名,但当前代码无法实现这个逻辑。现有代码如下:
import os import requests user_input = "wget google.com" parts = user_input.split() if parts[0] == "wget": # Check if the user just typed in wget if user_input == "wget": print( "wget: missing URL\nUsage: wget [URL]..." ) # If not, run wget else: try: # Store the url the user put in in a variable url = parts[1] # Get the file name from the url file_name = url.split("/")[-1] # Destination destination = os.path.join(os.getcwd(),file_name) # Checking if the user typed in a url with http or https. If they didn't, it will add http:// to the url. if not url.startswith("https://") and not url.startswith("http://"): url = "http://" + url # Send an HTTP GET request to the URL response = requests.get(url) # If the response status code is not 200, raise an exception response.raise_for_status() # Write the response content to the destination file with open(destination, 'wb') as f: f.write(response.content) print("'{}' saved".format(file_name)) # Catching any exception except Exception as e: print("wget: An error occurred:", e) print("Response status code:", response.status_code) print("Response content:", response.content)
解决方案
要实现区分网页和普通文件的命名逻辑,可以从响应的Content-Type头和URL路径特征两方面判断:
- 当URL以
/结尾(通常对应网站根目录或目录页),或者响应的Content-Type包含text/html时,判定为网页,使用index.html作为保存文件名; - 其他情况保留原URL提取的文件名。
修改后的代码如下(关键修改处已标注):
import os import requests user_input = "wget google.com" parts = user_input.split() if parts[0] == "wget": if user_input == "wget": print("wget: missing URL\nUsage: wget [URL]...") else: try: url = parts[1] # 初始从URL提取文件名 file_name = url.split("/")[-1] # 补全URL协议头 if not url.startswith("https://") and not url.startswith("http://"): url = "http://" + url response = requests.get(url) response.raise_for_status() # --- 新增:判断是否为网页,调整文件名 --- is_webpage = False # 1. 检查URL是否以/结尾(目录页) if url.endswith("/"): is_webpage = True # 2. 检查响应Content-Type是否为HTML content_type = response.headers.get("Content-Type", "") if "text/html" in content_type.lower(): is_webpage = True if is_webpage: file_name = "index.html" # --- 文件名调整结束 --- destination = os.path.join(os.getcwd(), file_name) with open(destination, 'wb') as f: f.write(response.content) print(f"'{file_name}' saved") except Exception as e: print("wget: An error occurred:", e) # 避免未定义response的情况 if 'response' in locals(): print("Response status code:", response.status_code) print("Response content:", response.content)
关键修改说明
- 在获取响应后,新增网页判断逻辑:结合URL结尾是否为
/和响应的Content-Type头判定是否为网页; - 若判定为网页,将文件名强制设为
index.html; - 异常处理中增加
response变量存在性判断,避免报错时因未定义response导致二次错误。
内容的提问来源于stack exchange,提问作者pizzavitdit
相关产品推荐
相关产品推荐

