HTTrack批量获取网站子页面独立HTML文件的方法咨询
Hey there! I’ve dealt with this exact HTTrack quirk before, so I totally get how frustrating it is to have the "Get separated files" option stop short after the main page. Let me share a few practical solutions to get all those subpages as individual HTML files without manually adding every URL:
Adjust your HTTrack mirror settings (the core fix)
When setting up a new project:- Enter your target URL and name the project as usual.
- Head to the Set options menu, then go to the Download tab.
- Make sure you’ve set a proper depth limit (or select "Mirror entire site" if that’s what you need) to capture all subpages.
- Here’s the key: Don’t use the "Get separated files" option right away. Stick with the default Mirror website structure setting. Once the mirror finishes, navigate to your project’s folder—you’ll find every subpage as a standalone .html file, organized in a directory structure matching the original site. The index.html is just a navigation hub; you can ignore it and access each individual file directly.
- Pro tip: If you want to filter out non-HTML files (like images, CSS), go to the Filters tab and add a rule like
+*.html +*.htmto only grab those file types.
Extract standalone files from an existing mirror
If you already have a completed mirror project, no need to re-scrape the site:
Just open your HTTrack project’s root folder. Dig into the site-specific subdirectory (or use your system’s search function to look for*.htmlfiles across the project). All your subpages are already there as individual files—you can just copy them to a new folder to separate them from the navigation index.Bonus: Batch convert HTML to PDF easily
Since you mentioned wanting to turn these files into PDFs later, here’s a quick trick:
Use a command-line tool likewkhtmltopdfto automate the process. For example, on Windows, create a batch file (convert.bat) in your HTML folder with this code:for %%f in (*.html) do wkhtmltopdf "%%f" "%%~nf.pdf"On Linux/macOS, use a shell script:
for file in *.html; do wkhtmltopdf "$file" "${file%.html}.pdf"; doneInstall wkhtmltopdf first, then run the script—it’ll convert every HTML file to a matching PDF in one go.
备注:内容来源于stack exchange,提问作者YL73

