You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析robots.txt并检测拼接URL的HTTP状态码求助

Hey there! Let's fix up your robots.txt parsing and URL checking code step by step. I spotted a couple of key issues that are keeping it from working as expected, plus some tweaks to make it more robust.

First, let's break down the problems in your current code:

  • When you loop with for x in result_data_set, you're actually iterating over the dictionary keys ("Disallowed" and "Allowed"), not the individual paths stored in each array. That's why your URL checks aren't targeting the right endpoints.
  • You have duplicate import os statements, plus some unused imports (io, urllib.parse) that can be cleaned up.
  • The path parsing doesn't handle comments or extra whitespace in the robots.txt file, which might lead to invalid paths being stored.
  • Your success message doesn't show the actual URL or status code, making it hard to verify results.

Here's the fixed and improved code:

import urllib.request
import urllib.error

# Uncomment this line to let users input their own URL
# url = input("Input Url:\n")
url = 'https://stackoverflow.com/robots.txt'
raw_robots = urllib.request.urlopen(url)
robots = raw_robots.read().decode('utf-8')

result_data_set = {"Disallowed": [], "Allowed": []}

for line in robots.split("\n"):
    # Clean up the line and skip empty lines/comments
    cleaned_line = line.strip()
    if not cleaned_line or cleaned_line.startswith('#'):
        continue
    
    if cleaned_line.startswith('Allow:'):
        # Extract path, ignoring comments after the path
        path = cleaned_line.split(': ', 1)[1].split('#')[0].strip()
        result_data_set["Allowed"].append(path)
    elif cleaned_line.startswith('Disallow:'):
        path = cleaned_line.split(': ', 1)[1].split('#')[0].strip()
        result_data_set["Disallowed"].append(path)

print("Parsed robots.txt rules:\n", result_data_set)

base_url = 'https://stackoverflow.com'

# Check Allowed paths
print("\n--- Checking Allowed Paths ---")
for path in result_data_set["Allowed"]:
    full_url = base_url + path
    try:
        response = urllib.request.urlopen(full_url)
        print(f"✅ {full_url} - Status Code: {response.getcode()}")
    except urllib.error.HTTPError as e:
        print(f"❌ {full_url} - HTTP Error: {e.code}")
    except urllib.error.URLError as e:
        print(f"⚠️ {full_url} - Connection Error: {e.reason}")

# Check Disallowed paths
print("\n--- Checking Disallowed Paths ---")
for path in result_data_set["Disallowed"]:
    full_url = base_url + path
    try:
        response = urllib.request.urlopen(full_url)
        print(f"✅ {full_url} - Status Code: {response.getcode()}")
    except urllib.error.HTTPError as e:
        print(f"❌ {full_url} - HTTP Error: {e.code}")
    except urllib.error.URLError as e:
        print(f"⚠️ {full_url} - Connection Error: {e.reason}")

Key improvements made:

  1. Cleaner imports: Removed unused and duplicate imports to keep things lean.
  2. Robust line processing: Skips empty lines and comments, and properly extracts paths even if they have trailing comments (like Allow: /abc # some note).
  3. Correct path iteration: Now loops directly over the paths in each array (Allowed and Disallowed) instead of dictionary keys.
  4. Clearer output: Shows the full URL, status code, and uses emojis to make results easy to scan at a glance.
  5. Separated checks: Groups results for allowed and disallowed paths to keep output organized.

A quick note: Some robots.txt files use wildcards (like /abc*) or relative paths that your current code won't handle, but this version fixes the core functionality you asked for.

内容的提问来源于stack exchange,提问作者Blank95

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:48:54