You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

wget镜像不跟随链接问题:半镜像站点下载MP3并避免重复下载

Hey there! Let's work through this wget issue you're facing—getting it to properly follow those "next page" links so you can grab all the MP3s on the site, without re-downloading files you already have. I've dealt with similar scenarios before, so let's break down the key options you might be missing.

First, Fix Recursion Depth & Scope

The most common reason wget stops at the first page is that it's hitting a default recursion limit or wandering outside your target site. Try these options:

  • -l inf: Sets infinite recursion depth. By default, wget only goes 5 levels deep, which might not be enough to reach all pages with MP3s.
  • --no-parent: Prevents wget from crawling up to parent directories outside your starting page's path—this keeps it focused on the content you care about.
  • --domain=your-target-site.com: Explicitly limits wget to only follow links within the target domain. This avoids accidental jumps to external sites and keeps the crawl focused.

Sometimes "next page" links might be in non-standard tags, or wget might be filtering out pages it doesn't think are relevant. Try these tweaks:

  • --follow-tags=a: Explicitly tells wget to follow links in <a> tags (this is default, but some sites use JavaScript-driven links that wget can't parse—more on that later).
  • --accept mp3,html: Makes wget download both HTML pages (to scan for MP3 links) and MP3 files. Without including HTML, wget won't have pages to parse for "next page" or MP3 links.
  • --adjust-extension: If some MP3 links don't end with .mp3, this adds the correct extension to downloaded files, making it easier to organize and avoid duplicates.

Avoid Duplicates & Handle Interruptions

You already used -N -c, but let's reinforce those and add a safety net:

  • -N: Checks the server's timestamp against your local files—only downloads if the server's file is newer, so no redundant downloads.
  • -c: Resumes interrupted downloads, which is helpful if your crawl gets cut off.
  • --no-clobber: Prevents wget from overwriting existing files (though -N should handle this, it's an extra layer of protection).

Beat Anti-Crawl Measures

Many sites block wget by default or limit frequent requests. Add these to avoid being blocked:

  • --user-agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36": Makes wget mimic a browser, so the server doesn't reject it.
  • --wait=2 --random-wait: Adds a 2-second delay between requests, with random variation. This prevents triggering rate limits that might block your crawl.

Example Command

Putting it all together, here's a command that should handle most cases:

wget -r -l inf -N -c --accept mp3,html --no-parent --domain=your-target-site.com --user-agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" --wait=2 --random-wait your-target-site.com/starting-page/

If the "next page" links are loaded via JavaScript (not standard <a> tags), wget won't be able to parse them directly. In that case, you might need a tool like yt-dlp with recursive options, or use a headless browser to scrape links first, then feed them to wget. But since you asked for wget options, the above should cover most static link scenarios.

内容的提问来源于stack exchange,提问作者Edward

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:20:30