如何检查文件中URL列表有效性?Bash脚本新手批量检测遇阻求助
Hey there! Let's troubleshoot why your curl batch URL check is incorrectly shoving active URLs into nv.txt—this is such a common snag for Bash newbies, so let's break down the fixes step by step.
First, Let's Diagnose the Common Issues
From your description, here are the most likely culprits:
- Incorrect URL reading: Using
catwithwhile readcan mangle URLs with spaces, backslashes, or leading/trailing whitespace. - Misinterpreting curl's exit code: By default, curl returns a successful exit code (0) even if the server sends a 404/500 error—you might be marking valid-but-erroring URLs as inactive.
- Ignoring redirects: Many active URLs use 301/302 redirects; if curl doesn't follow them, you'll incorrectly flag those as inactive.
- No timeout handling: Slow/unresponsive URLs can hang your script or cause false negatives.
Fixed Bash Script
Here's a revised script that addresses all these issues, with comments explaining each part:
#!/bin/bash # Configure your files here URL_LIST="urls.txt" # Your input list of URLs ACTIVE_OUTPUT="active.txt"# File for working URLs INACTIVE_OUTPUT="nv.txt" # File for non-working URLs # Clear output files (optional: removes old results before starting) > "$ACTIVE_OUTPUT" > "$INACTIVE_OUTPUT" # Properly read each URL (avoids mangling special characters/whitespace) while IFS= read -r url; do # Skip empty lines in your URL list [[ -z "$url" ]] && continue # Run curl with safe, reliable parameters # -s: Silent mode (no progress bars) # -L: Follow redirects (critical for URLs that forward to other pages) # -f: Return non-zero exit code for 4xx/5xx errors (marks broken-but-reachable URLs as inactive) # --max-time 10: Time out after 10 seconds (prevents hanging) # -o /dev/null: Discard page content (we don't need it) # -w "%{http_code}": Capture the HTTP status code for debugging http_status=$(curl -s -L -f --max-time 10 -o /dev/null -w "%{http_code}" "$url") curl_exit_code=$? # Classify the URL based on curl's exit code if [[ $curl_exit_code -eq 0 ]]; then echo "$url (HTTP Status: $http_status)" >> "$ACTIVE_OUTPUT" echo "✅ Active: $url" # Optional: Print progress to terminal else echo "$url (Exit Code: $curl_exit_code, HTTP Status: $http_status)" >> "$INACTIVE_OUTPUT" echo "❌ Inactive: $url" # Optional: Print progress to terminal fi done < "$URL_LIST" echo -e "\nDone! Results saved:" echo "- Active URLs: $ACTIVE_OUTPUT" echo "- Inactive URLs: $INACTIVE_OUTPUT"
Key Fixes Explained
Safe URL Reading:
while IFS= read -r urlis the standard way to read lines in Bash:IFS=preserves leading/trailing whitespace in URLs-rdisables backslash escaping, so URLs with\aren't corrupted- Reading directly from
$URL_LIST(instead of pipingcat) avoids subshell issues (like variables not persisting outside the loop)
Curl Parameter Tweaks:
-Lensures we follow redirects (so URLs likehttp://example.comthat forward tohttps://www.example.comaren't marked inactive)-fchanges curl's behavior to return an error code for 4xx/5xx responses—if you consider "active" as any reachable URL (even if it returns 404), remove this flag--max-time 10prevents your script from getting stuck on slow or unresponsive servers
Debugging Context:
The script captures both curl's exit code and the HTTP status code, so you can see why a URL was classified (e.g., a 503 means the server is down temporarily, while exit code 7 means curl couldn't connect at all).
How to Use
- Save your list of URLs in a file named
urls.txt(one URL per line) - Save the script as
check_urls.sh - Make it executable:
chmod +x check_urls.sh - Run it:
./check_urls.sh
内容的提问来源于stack exchange,提问作者Jonathan Mann

