You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDF批量下载脚本仅部分下载且速度慢的问题排查与优化

Fixing the Leipziger Zeitung Download Script: Full Issue Coverage & Speed Boosts

Let's get your script to download every available issue of the Leipziger Zeitung (1814-1857) and cut down on runtime. Here's a revised version with explanations of the key fixes:

Revised Script

#!/bin/bash
# Script to download ALL issues of the Leipziger Zeitung (1814-1857) with parallel processing

# Mimic a real browser to avoid server blocks (keep this updated if needed)
USER_AGENT="Mozilla/5.0 (Macintosh; Intel Mac OS X 10_13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/66.0.3359.139 Safari/537.36"

# Loop through each year (1814 to 1857)
for year in {14..57}; do
  echo "Processing year 18$year..."
  
  # Extract ALL unique 8-digit date values from the year's page
  DATES=$(curl -sS "http://anno.onb.ac.at/cgi-content/anno?aid=lzg&datum=18$year" | 
          grep -oP 'datum=\K\d{8}' | sort -u)
  
  # Skip if no dates were found for the year
  if [ -z "$DATES" ]; then
    echo "No issues found for 18$year, moving on..."
    continue
  fi
  
  # Download PDFs in parallel (adjust -P to 6-8 if your network can handle it)
  echo "$DATES" | xargs -P 4 -I {} bash -c '
    DATE="{}"
    echo "Downloading: ONB_lzg_${DATE}.pdf"
    wget -nc -E -nd --no-check-certificate --content-disposition \
         --user-agent="$USER_AGENT" \
         --header="Referer: http://anno.onb.ac.at/cgi-content/anno?aid=lzg&datum=${DATE}" \
         "http://anno.onb.ac.at/pdfs/ONB_lzg_${DATE}.pdf"
  '
done

What Changed & Why

1. Fixing Missing Issues

The original script’s gawk regex only captured a subset of dates because it didn’t target the full 8-digit date format used on the site. Here's the fix:

  • Accurate Date Extraction: grep -oP 'datum=\K\d{8}' specifically pulls out every datum= parameter followed by an 8-digit date (YYYYMMDD), ensuring we catch every available issue.
  • Deduplication: sort -u removes duplicate date entries that might appear in the page’s HTML, so we don’t waste time re-downloading files.

2. Speeding Up Downloads

Your original script ran downloads one at a time and included redundant curl calls—here’s how we fixed that:

  • Parallel Processing: xargs -P 4 runs up to 4 download processes simultaneously (tweak the number based on your network speed; 4-8 is safe). This cuts total runtime drastically compared to serial downloads.
  • Removed Redundant curl: The original curl request for each date was unnecessary. We moved all required headers (User-Agent, Referer) directly into the wget command, eliminating that 3-5 second delay per issue.
  • Faster Date Parsing: Using grep instead of gawk for pattern matching is more efficient for this specific task.

3. Added Robustness

  • Error Checking: The script now skips years where no dates are found (though this shouldn’t happen for 1814-1857) instead of failing silently.
  • Browser Mimicry: We kept the original User-Agent and Referer headers to ensure the server accepts our requests—mimicking a real browser helps avoid being blocked.

内容的提问来源于stack exchange,提问作者Til Hund

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:02:21