PDF批量下载脚本仅部分下载且速度慢的问题排查与优化
Fixing the Leipziger Zeitung Download Script: Full Issue Coverage & Speed Boosts
Let's get your script to download every available issue of the Leipziger Zeitung (1814-1857) and cut down on runtime. Here's a revised version with explanations of the key fixes:
Revised Script
#!/bin/bash # Script to download ALL issues of the Leipziger Zeitung (1814-1857) with parallel processing # Mimic a real browser to avoid server blocks (keep this updated if needed) USER_AGENT="Mozilla/5.0 (Macintosh; Intel Mac OS X 10_13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/66.0.3359.139 Safari/537.36" # Loop through each year (1814 to 1857) for year in {14..57}; do echo "Processing year 18$year..." # Extract ALL unique 8-digit date values from the year's page DATES=$(curl -sS "http://anno.onb.ac.at/cgi-content/anno?aid=lzg&datum=18$year" | grep -oP 'datum=\K\d{8}' | sort -u) # Skip if no dates were found for the year if [ -z "$DATES" ]; then echo "No issues found for 18$year, moving on..." continue fi # Download PDFs in parallel (adjust -P to 6-8 if your network can handle it) echo "$DATES" | xargs -P 4 -I {} bash -c ' DATE="{}" echo "Downloading: ONB_lzg_${DATE}.pdf" wget -nc -E -nd --no-check-certificate --content-disposition \ --user-agent="$USER_AGENT" \ --header="Referer: http://anno.onb.ac.at/cgi-content/anno?aid=lzg&datum=${DATE}" \ "http://anno.onb.ac.at/pdfs/ONB_lzg_${DATE}.pdf" ' done
What Changed & Why
1. Fixing Missing Issues
The original script’s gawk regex only captured a subset of dates because it didn’t target the full 8-digit date format used on the site. Here's the fix:
- Accurate Date Extraction:
grep -oP 'datum=\K\d{8}'specifically pulls out everydatum=parameter followed by an 8-digit date (YYYYMMDD), ensuring we catch every available issue. - Deduplication:
sort -uremoves duplicate date entries that might appear in the page’s HTML, so we don’t waste time re-downloading files.
2. Speeding Up Downloads
Your original script ran downloads one at a time and included redundant curl calls—here’s how we fixed that:
- Parallel Processing:
xargs -P 4runs up to 4 download processes simultaneously (tweak the number based on your network speed; 4-8 is safe). This cuts total runtime drastically compared to serial downloads. - Removed Redundant curl: The original curl request for each date was unnecessary. We moved all required headers (User-Agent, Referer) directly into the wget command, eliminating that 3-5 second delay per issue.
- Faster Date Parsing: Using
grepinstead ofgawkfor pattern matching is more efficient for this specific task.
3. Added Robustness
- Error Checking: The script now skips years where no dates are found (though this shouldn’t happen for 1814-1857) instead of failing silently.
- Browser Mimicry: We kept the original User-Agent and Referer headers to ensure the server accepts our requests—mimicking a real browser helps avoid being blocked.
内容的提问来源于stack exchange,提问作者Til Hund
相关产品推荐
相关产品推荐

