使用Wget递归下载网站时无法检测SVG文件的问题求助
Hey, let's work through why Wget isn't picking up those SVG files hosted on the assets subdomain! I've run into similar issues before, so here are some troubleshooting steps and fixes to try:
First, check how your SVGs are referenced
Wget's --page-requisites flag works great for assets linked directly in HTML tags (like <img src="...svg">), but it might miss SVGs loaded in other ways:
- Via CSS background images (e.g.,
background: url(/assets/icon.svg)) - Dynamically injected with JavaScript
- Embedded using
<object>or<embed>tags
If your SVGs are in CSS or JS, Wget won't automatically parse those files to extract links by default.
Adjust your Wget command to explicitly target SVGs
Let's tweak your command to make sure Wget prioritizes and detects SVG files. Add the --accept flag to whitelist SVG (along with other critical file types) and double-check your domain settings:
wget --load-cookies cookies.txt --mirror --recursive --level=5 --no-parent --page-requisites --domains=LINK,assets.LINK -H --convert-links --adjust-extension --header="referer: LINK" --accept=html,css,js,svg LINK
A couple key notes here:
- The
--accept=html,css,js,svgtells Wget to only grab these file types (you can add others if needed) - Double-check that
assets.LINKmatches the actual subdomain hosting your SVGs—typos here are a common culprit!
Debug with Wget's verbose logging
To see exactly why Wget is skipping the SVGs, add the --debug flag to your command. This will spit out detailed logs that show:
- If Wget even encounters the SVG links in the first place
- If the links are being rejected because they're outside your specified
--domains - If you're getting 403 Forbidden errors (which might mean your referer/cookie settings need adjustment)
Look for lines like Rejected because not in any domain (fix by updating --domains) or HTTP request sent, awaiting response... 403 Forbidden (check your referer header or cookie validity).
Handle SVGs embedded in CSS
If your SVGs are referenced in CSS files, Wget won't parse CSS to extract those links on its own. Here's a simple workaround:
- First, run your original Wget command to download all HTML and CSS files
- Extract the SVG links from your CSS files using a grep command:
grep -o 'url(["'\''"]\?https\?://assets\.LINK/[^"'\''")]\+\.svg["'\''"]\?' *.css | cut -d'(' -f2 | tr -d '"\')' > svg_urls.txt - Use Wget to batch-download those extracted SVG links:
Thewget --load-cookies cookies.txt --header="referer: LINK" -i svg_urls.txt -P assets.LINK/-Pflag ensures the SVGs are saved to the correct subdomain directory to match your original site structure.
Check for conflicts with --mirror
The --mirror flag enables several options automatically, including --level=inf (infinite recursion). Since you're specifying --level=5 manually, there might be an unexpected conflict. Try removing --mirror and using explicit flags instead:
wget --load-cookies cookies.txt --recursive --level=5 --timestamping --no-parent --page-requisites --domains=LINK,assets.LINK -H --convert-links --adjust-extension --header="referer: LINK" --accept=svg,html,css,js LINK
Give these steps a shot—chances are one of them will get those SVGs downloading properly!
备注:内容来源于stack exchange,提问作者ChickenHead

