Bash脚本循环意外终止求助:批量获取服务器信息中途停止
Hey Philip, let's dig into why your script is dying around the 45th URL—this is a super common pain point when looping through lots of curl calls, so let's break down the likely culprits and fixes.
First, Let's Cover the Most Probable Issues
From the code snippets you shared, here are the top things to check:
1. Unhandled curl Failures Stopping the Script
Your curl command uses --fail, which makes curl exit with a non-zero code if the HTTP response isn't a 200-series success. By default, Bash stops executing the entire script when a command returns a non-zero exit code (unless you've disabled this with set +e). If the 45th URL returns a 404, 500, or any non-success status, your script will grind to a halt.
Fix: Add error handling to let the loop continue even if a single curl call fails. For example:
# Wrap curl in an if/else to catch failures if curl $curlOptions -D "$headerFile" -o "$dataPath/$(basename $url)" "$dataURL$url"; then echo "Success: $url" >> script_progress.log else echo "Failed (exit code $?): $url" >> script_errors.log continue # Skip to the next URL instead of stopping fi
2. Missing Full Request Timeout
You set timeout=5 and used --connect-timeout $time... (note: looks like a typo here—should be $timeout!). But --connect-timeout only applies to the initial connection phase. If a URL connects successfully but takes forever to send data, your script will hang indefinitely.
Fix: Add --max-time $timeout to your curl options to set a hard timeout for the entire request:
curlOptions="--fail --connect-timeout $timeout --max-time $timeout -s"
3. Broken URL Reading from the TSV File
If you're using a basic for loop to read URLs from fortune500.tsv, you'll run into issues with URLs containing special characters (like & or spaces) or TSV fields splitting incorrectly.
Fix: Use while read with proper TSV handling to safely parse each line:
# Assuming the first column in your TSV is the URL while IFS=$'\t' read -r url other_fields; do # Your curl logic here done < "$dataFile"
The -r flag prevents backslashes in URLs from being interpreted as escape characters, and IFS=$'\t' tells Bash to split fields on tabs instead of spaces.
4. Unchecked File/Directory Permissions
If $dataPath doesn't exist, or your user doesn't have write permissions to $headerFile or the data directory, the script will stop when it tries to write to a location it can't access.
Fix: Add a check to create the data directory if it doesn't exist:
mkdir -p "$dataPath"
Debugging Step to Pinpoint the Exact Issue
Before modifying the script, run this quick test to see what's happening with the 45th URL:
- Find the 45th line in your TSV:
sed -n '45p' fortune500.tsv - Copy that URL and run curl manually with your options:
curl --fail --connect-timeout 5 --max-time 5 -D test_header.txt -o test_output.txt "http://www.tech.mtu.edu/~toarney/sat3310/lab09/[YOUR_45TH_URL]"
This will show you exactly why curl is failing (e.g., 404 error, connection timeout, etc.), which will guide your fix.
Example Fixed Loop Snippet
Putting it all together, your loop might look like this:
# Initialize variables timeout=5 headerFile="lab06.output" dataFile="fortune500.tsv" dataURL="http://www.tech.mtu.edu/~toarney/sat3310/lab09/" dataPath="/home/pjvaglic/Documents/labs/lab06/data/" curlOptions="--fail --connect-timeout $timeout --max-time $timeout -s" # Ensure data directory exists mkdir -p "$dataPath" # Track progress count=0 # Read URLs from TSV safely while IFS=$'\t' read -r url unused_fields; do count=$((count + 1)) echo "Processing URL $count/500: $url" >> script_progress.log # Attempt curl with error handling if curl $curlOptions -D "$headerFile" -o "$dataPath/$(basename $url)" "$dataURL$url"; then echo "Successfully fetched $url" >> script_progress.log else echo "Failed to fetch $url (exit code: $?) at $(date)" >> script_errors.log continue fi done < "$dataFile"
内容的提问来源于stack exchange,提问作者Philip

