如何用wget克隆网站指定目录及子目录?现有命令仅获index.html
解决wget仅克隆index.html的问题及替代Python实现
一、wget命令的问题修复
你当前的wget命令缺少**--no-parent(缩写-np)**参数,这是导致只拿到index.html的核心原因:
- 默认情况下,wget递归爬取时会尝试访问父目录的内容,但很多网站的目录结构会限制跨目录访问,或者wget无法正确识别子目录入口,加上
-np可以强制wget只爬取指定目录及其子目录的内容,不会向上回溯父目录。
修正后的完整命令:
wget --limit-rate=700k --no-clobber --convert-links --random-wait -r -p -E -e robots=off -U mozilla --no-parent https://www.offsec.com/metasploit-unleashed/
补充说明:
- 你已保留URL末尾的斜杠
/,这是正确的,避免wget把metasploit-unleashed当成单个文件处理。 -p参数会下载页面所需的所有资源(css、js、图片等),配合--convert-links能确保离线打开时资源路径正确。
二、Python实现替代方案
如果wget仍然有问题,可以用Python脚本实现定向爬取,以下代码支持递归爬取目标目录、资源下载、本地链接转换和限速:
import os import time import requests from bs4 import BeautifulSoup from urllib.parse import urljoin, urlparse # 配置参数 BASE_URL = "https://www.offsec.com/metasploit-unleashed/" LOCAL_SAVE_DIR = "metasploit-unleashed-offline" USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" MAX_DOWNLOAD_SPEED = 700 * 1024 # 700KB/s,换算为字节/秒 # 创建本地保存目录 os.makedirs(LOCAL_SAVE_DIR, exist_ok=True) def is_in_target_scope(url): """判断链接是否属于目标目录的范围""" parsed_target = urlparse(BASE_URL) parsed_url = urlparse(url) # 检查域名一致,且路径以目标目录开头 return parsed_url.netloc == parsed_target.netloc and parsed_url.path.startswith(parsed_target.path) def download_with_speed_limit(url, save_path): """限速下载文件""" try: response = requests.get(url, headers={"User-Agent": USER_AGENT}, stream=True) response.raise_for_status() os.makedirs(os.path.dirname(save_path), exist_ok=True) start_time = time.time() downloaded_bytes = 0 with open(save_path, 'wb') as f: for chunk in response.iter_content(chunk_size=1024): if chunk: f.write(chunk) downloaded_bytes += len(chunk) # 计算已用时间和理论应下载字节数,控制速度 elapsed_time = time.time() - start_time expected_bytes = MAX_DOWNLOAD_SPEED * elapsed_time if downloaded_bytes > expected_bytes: time.sleep((downloaded_bytes - expected_bytes) / MAX_DOWNLOAD_SPEED) except Exception as e: print(f"下载失败 {url}: {str(e)}") def crawl_and_save_page(url, visited_urls): """递归爬取页面并保存为本地文件""" if url in visited_urls: return visited_urls.add(url) try: # 获取页面内容 response = requests.get(url, headers={"User-Agent": USER_AGENT}) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') # 生成本地保存路径 parsed_url = urlparse(url) local_file_path = os.path.join(LOCAL_SAVE_DIR, parsed_url.path.lstrip('/')) # 处理目录形式的URL,默认保存为index.html if not local_file_path.endswith(('.html', '.htm')): if '.' not in os.path.basename(local_file_path): local_file_path = os.path.join(local_file_path, 'index.html') os.makedirs(os.path.dirname(local_file_path), exist_ok=True) # 遍历并转换页面中的所有资源链接 for tag in soup.find_all(['a', 'link', 'script', 'img', 'source']): attr = 'href' if tag.name in ['a', 'link'] else 'src' if not tag.has_attr(attr): continue original_link = tag[attr] absolute_link = urljoin(url, original_link) if is_in_target_scope(absolute_link): # 生成资源的本地路径 parsed_link = urlparse(absolute_link) local_resource_path = os.path.join(LOCAL_SAVE_DIR, parsed_link.path.lstrip('/')) # 转换为相对本地路径 relative_path = os.path.relpath(local_resource_path, os.path.dirname(local_file_path)) tag[attr] = relative_path # 判断是页面还是静态资源,页面则递归爬取,资源直接下载 if absolute_link.endswith('/') or '.' not in os.path.basename(parsed_link.path): crawl_and_save_page(absolute_link, visited_urls) else: download_with_speed_limit(absolute_link, local_resource_path) # 保存处理后的页面 with open(local_file_path, 'w', encoding='utf-8') as f: f.write(str(soup)) print(f"已保存: {local_file_path}") except Exception as e: print(f"爬取失败 {url}: {str(e)}") # 启动爬取 visited = set() crawl_and_save_page(BASE_URL, visited) print(f"离线克隆完成,所有文件已保存至: {os.path.abspath(LOCAL_SAVE_DIR)}")
使用说明:
- 确保已安装依赖库:
pip install requests beautifulsoup4 - 运行脚本后,所有页面和资源会保存到当前目录下的
metasploit-unleashed-offline文件夹 - 脚本自带限速功能,模拟wget的
--limit-rate效果,同时自动转换本地链接,保证离线预览正常
内容的提问来源于stack exchange,提问作者user14737770
相关产品推荐
相关产品推荐

