为何wget --reject-regex .*命令仍能从www.example.com下载index.html?
wget --reject-regex .* http://www.example.com/ still downloads index.html? This is a super common gotcha with wget's --reject-regex flag—let's unpack exactly what's going on here:
--reject-regexfilters discovered links, not your initial target URL
When you run wget againsthttp://www.example.com/, wget recognizes this as a directory path and automatically requests the default index file (typicallyindex.html) for that directory. This initial request for the index file is not evaluated against your--reject-regexrule. The regex only applies to links that wget finds after fetching the initial page (think: links to other pages, images, CSS files, etc.).Your regex
.*is correct, but it never gets applied to the initial index.html request
The.*pattern absolutely would match any full URL if it were checked. But wget's design prioritizes fetching the starting URL you specify—even if it technically matches a reject filter. Reject rules are meant to trim the crawl, not block the entry point.
How to confirm this behavior
Add the --debug flag to your command to see the exact requests wget makes and how it applies filters:
wget --debug --reject-regex .* http://www.example.com/
In the output, look for lines referencing index.html—you'll notice wget never runs the reject regex against this initial request.
If you want to block all downloads (including the initial index.html)
If your goal is to stop wget from downloading anything at all, here are two clean options:
- Use the
--spiderflag: This turns wget into a spider that only checks if the URL is reachable, with no local file downloads. - Combine
--reject-regex .*with--no-directoriesand--delete-after, but--spideris far simpler for this use case. Alternatively, if you specifically want to block the index file, add--reject="index.html"to your command (though combining it with.*is redundant for full blocking).
内容的提问来源于stack exchange,提问作者Brian

