You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置HTTrack仅下载指定子域名内容?Scan Rules设置无效求助

Fixing HTTrack to Download Full Slader Textbook Content

Hey there, let's get your HTTrack setup sorted so you can save all that Calculus textbook content from Slader—including nested subpages (like /311/) and embedded images—no more 10-second incomplete downloads! 🛠️

Step 1: Nail Down Basic Configuration

First, let's fix the core settings that are likely causing the partial download:

  • Double-check your starting URL: Make sure you've entered https://www.slader.com/textbook/9781337624183-calculus-9th-edition/ as the primary URL to mirror.
  • Adjust mirror depth: Head to the "Set options" tab → "Limits". Set the mirror depth to at least 5 (textbook content is nested deeply, so a shallow depth only grabs the main page). Avoid setting it to 0 or 1!
  • Enable media and embedded content: Still in "Set options", check the boxes for "Download images" and "Download embedded objects (scripts, stylesheets, etc.)". This ensures those embedded images (even from other domains) get saved alongside pages.

Step 2: Configure Scan Rules Correctly

This is where most people get tripped up—HTTrack's rules need to prioritize what to allow before blocking other content. Here's the exact rule set to add in the "Scan Rules" tab (order matters!):

  1. Allow all subpages under your textbook:
    +www.slader.com/textbook/9781337624183-calculus-9th-edition/*
    
    This tells HTTrack to follow every link nested under your specific textbook page.
  2. Allow all common image formats (to grab embedded images from any domain):
    +*.jpg +*.jpeg +*.png +*.gif +*.svg
    
  3. Block all other external links (to avoid wasting time on unrelated content):
    -*
    
    By putting allow rules first, HTTrack will only ignore content that doesn't match your textbook or image rules.

Step 3: Extra Tweaks to Catch All Content

A few more settings to ensure nothing slips through:

  • Turn off "Stay on same domain": In "Set options" → "Limits", uncheck this box. This lets HTTrack grab images hosted on other domains (which Slader often uses) while still following your scan rules to stay focused.
  • Enable dynamic link following: Under "Set options" → "Browser emulation", check "Follow links in scripts/styles". Slader loads some content dynamically, so this helps HTTrack catch links hidden in JS or CSS files.
  • Increase timeout: In "Set options" → "Connection", set the connection timeout to 30-60 seconds. Slader can be slow to load pages, and a short timeout might make HTTrack give up before grabbing content.

Step 4: Test Before Full Mirror

Before running a full download, do a quick test: set the mirror depth to 2, start the download, and check if the main page plus a couple of subpages (like /311/) are saved correctly. If that works, you can bump up the depth and run the full mirror.

内容的提问来源于stack exchange,提问作者Unrivalled confusion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 00:37:32