You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用wget克隆网站指定目录及子目录?现有命令仅获index.html

解决wget仅克隆index.html的问题及替代Python实现

一、wget命令的问题修复

你当前的wget命令缺少**--no-parent(缩写-np)**参数,这是导致只拿到index.html的核心原因:

  • 默认情况下,wget递归爬取时会尝试访问父目录的内容,但很多网站的目录结构会限制跨目录访问,或者wget无法正确识别子目录入口,加上-np可以强制wget只爬取指定目录及其子目录的内容,不会向上回溯父目录。

修正后的完整命令:

wget --limit-rate=700k --no-clobber --convert-links --random-wait -r -p -E -e robots=off -U mozilla --no-parent https://www.offsec.com/metasploit-unleashed/

补充说明:

  • 你已保留URL末尾的斜杠/,这是正确的,避免wget把metasploit-unleashed当成单个文件处理。
  • -p参数会下载页面所需的所有资源(css、js、图片等),配合--convert-links能确保离线打开时资源路径正确。

二、Python实现替代方案

如果wget仍然有问题,可以用Python脚本实现定向爬取,以下代码支持递归爬取目标目录、资源下载、本地链接转换和限速:

import os
import time
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse

# 配置参数
BASE_URL = "https://www.offsec.com/metasploit-unleashed/"
LOCAL_SAVE_DIR = "metasploit-unleashed-offline"
USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
MAX_DOWNLOAD_SPEED = 700 * 1024  # 700KB/s,换算为字节/秒

# 创建本地保存目录
os.makedirs(LOCAL_SAVE_DIR, exist_ok=True)

def is_in_target_scope(url):
    """判断链接是否属于目标目录的范围"""
    parsed_target = urlparse(BASE_URL)
    parsed_url = urlparse(url)
    # 检查域名一致,且路径以目标目录开头
    return parsed_url.netloc == parsed_target.netloc and parsed_url.path.startswith(parsed_target.path)

def download_with_speed_limit(url, save_path):
    """限速下载文件"""
    try:
        response = requests.get(url, headers={"User-Agent": USER_AGENT}, stream=True)
        response.raise_for_status()
        
        os.makedirs(os.path.dirname(save_path), exist_ok=True)
        start_time = time.time()
        downloaded_bytes = 0
        
        with open(save_path, 'wb') as f:
            for chunk in response.iter_content(chunk_size=1024):
                if chunk:
                    f.write(chunk)
                    downloaded_bytes += len(chunk)
                    # 计算已用时间和理论应下载字节数,控制速度
                    elapsed_time = time.time() - start_time
                    expected_bytes = MAX_DOWNLOAD_SPEED * elapsed_time
                    if downloaded_bytes > expected_bytes:
                        time.sleep((downloaded_bytes - expected_bytes) / MAX_DOWNLOAD_SPEED)
    except Exception as e:
        print(f"下载失败 {url}: {str(e)}")

def crawl_and_save_page(url, visited_urls):
    """递归爬取页面并保存为本地文件"""
    if url in visited_urls:
        return
    visited_urls.add(url)
    
    try:
        # 获取页面内容
        response = requests.get(url, headers={"User-Agent": USER_AGENT})
        response.raise_for_status()
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # 生成本地保存路径
        parsed_url = urlparse(url)
        local_file_path = os.path.join(LOCAL_SAVE_DIR, parsed_url.path.lstrip('/'))
        # 处理目录形式的URL,默认保存为index.html
        if not local_file_path.endswith(('.html', '.htm')):
            if '.' not in os.path.basename(local_file_path):
                local_file_path = os.path.join(local_file_path, 'index.html')
        
        os.makedirs(os.path.dirname(local_file_path), exist_ok=True)
        
        # 遍历并转换页面中的所有资源链接
        for tag in soup.find_all(['a', 'link', 'script', 'img', 'source']):
            attr = 'href' if tag.name in ['a', 'link'] else 'src'
            if not tag.has_attr(attr):
                continue
            
            original_link = tag[attr]
            absolute_link = urljoin(url, original_link)
            
            if is_in_target_scope(absolute_link):
                # 生成资源的本地路径
                parsed_link = urlparse(absolute_link)
                local_resource_path = os.path.join(LOCAL_SAVE_DIR, parsed_link.path.lstrip('/'))
                # 转换为相对本地路径
                relative_path = os.path.relpath(local_resource_path, os.path.dirname(local_file_path))
                tag[attr] = relative_path
                
                # 判断是页面还是静态资源,页面则递归爬取,资源直接下载
                if absolute_link.endswith('/') or '.' not in os.path.basename(parsed_link.path):
                    crawl_and_save_page(absolute_link, visited_urls)
                else:
                    download_with_speed_limit(absolute_link, local_resource_path)
        
        # 保存处理后的页面
        with open(local_file_path, 'w', encoding='utf-8') as f:
            f.write(str(soup))
        print(f"已保存: {local_file_path}")
    
    except Exception as e:
        print(f"爬取失败 {url}: {str(e)}")

# 启动爬取
visited = set()
crawl_and_save_page(BASE_URL, visited)
print(f"离线克隆完成,所有文件已保存至: {os.path.abspath(LOCAL_SAVE_DIR)}")

使用说明:

  1. 确保已安装依赖库:pip install requests beautifulsoup4
  2. 运行脚本后,所有页面和资源会保存到当前目录下的metasploit-unleashed-offline文件夹
  3. 脚本自带限速功能,模拟wget的--limit-rate效果,同时自动转换本地链接,保证离线预览正常

内容的提问来源于stack exchange,提问作者user14737770

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 03:35:40