You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何检查文件中URL列表有效性?Bash脚本新手批量检测遇阻求助

Hey there! Let's troubleshoot why your curl batch URL check is incorrectly shoving active URLs into nv.txt—this is such a common snag for Bash newbies, so let's break down the fixes step by step.

First, Let's Diagnose the Common Issues

From your description, here are the most likely culprits:

  • Incorrect URL reading: Using cat with while read can mangle URLs with spaces, backslashes, or leading/trailing whitespace.
  • Misinterpreting curl's exit code: By default, curl returns a successful exit code (0) even if the server sends a 404/500 error—you might be marking valid-but-erroring URLs as inactive.
  • Ignoring redirects: Many active URLs use 301/302 redirects; if curl doesn't follow them, you'll incorrectly flag those as inactive.
  • No timeout handling: Slow/unresponsive URLs can hang your script or cause false negatives.

Fixed Bash Script

Here's a revised script that addresses all these issues, with comments explaining each part:

#!/bin/bash

# Configure your files here
URL_LIST="urls.txt"       # Your input list of URLs
ACTIVE_OUTPUT="active.txt"# File for working URLs
INACTIVE_OUTPUT="nv.txt"  # File for non-working URLs

# Clear output files (optional: removes old results before starting)
> "$ACTIVE_OUTPUT"
> "$INACTIVE_OUTPUT"

# Properly read each URL (avoids mangling special characters/whitespace)
while IFS= read -r url; do
    # Skip empty lines in your URL list
    [[ -z "$url" ]] && continue

    # Run curl with safe, reliable parameters
    # -s: Silent mode (no progress bars)
    # -L: Follow redirects (critical for URLs that forward to other pages)
    # -f: Return non-zero exit code for 4xx/5xx errors (marks broken-but-reachable URLs as inactive)
    # --max-time 10: Time out after 10 seconds (prevents hanging)
    # -o /dev/null: Discard page content (we don't need it)
    # -w "%{http_code}": Capture the HTTP status code for debugging
    http_status=$(curl -s -L -f --max-time 10 -o /dev/null -w "%{http_code}" "$url")
    curl_exit_code=$?

    # Classify the URL based on curl's exit code
    if [[ $curl_exit_code -eq 0 ]]; then
        echo "$url (HTTP Status: $http_status)" >> "$ACTIVE_OUTPUT"
        echo "✅ Active: $url"  # Optional: Print progress to terminal
    else
        echo "$url (Exit Code: $curl_exit_code, HTTP Status: $http_status)" >> "$INACTIVE_OUTPUT"
        echo "❌ Inactive: $url"  # Optional: Print progress to terminal
    fi
done < "$URL_LIST"

echo -e "\nDone! Results saved:"
echo "- Active URLs: $ACTIVE_OUTPUT"
echo "- Inactive URLs: $INACTIVE_OUTPUT"

Key Fixes Explained

  1. Safe URL Reading:
    while IFS= read -r url is the standard way to read lines in Bash:

    • IFS= preserves leading/trailing whitespace in URLs
    • -r disables backslash escaping, so URLs with \ aren't corrupted
    • Reading directly from $URL_LIST (instead of piping cat) avoids subshell issues (like variables not persisting outside the loop)
  2. Curl Parameter Tweaks:

    • -L ensures we follow redirects (so URLs like http://example.com that forward to https://www.example.com aren't marked inactive)
    • -f changes curl's behavior to return an error code for 4xx/5xx responses—if you consider "active" as any reachable URL (even if it returns 404), remove this flag
    • --max-time 10 prevents your script from getting stuck on slow or unresponsive servers
  3. Debugging Context:
    The script captures both curl's exit code and the HTTP status code, so you can see why a URL was classified (e.g., a 503 means the server is down temporarily, while exit code 7 means curl couldn't connect at all).

How to Use

  1. Save your list of URLs in a file named urls.txt (one URL per line)
  2. Save the script as check_urls.sh
  3. Make it executable: chmod +x check_urls.sh
  4. Run it: ./check_urls.sh

内容的提问来源于stack exchange,提问作者Jonathan Mann

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:11:01