Python解析robots.txt并检测拼接URL的HTTP状态码求助
Hey there! Let's fix up your robots.txt parsing and URL checking code step by step. I spotted a couple of key issues that are keeping it from working as expected, plus some tweaks to make it more robust.
First, let's break down the problems in your current code:
- When you loop with
for x in result_data_set, you're actually iterating over the dictionary keys ("Disallowed"and"Allowed"), not the individual paths stored in each array. That's why your URL checks aren't targeting the right endpoints. - You have duplicate
import osstatements, plus some unused imports (io,urllib.parse) that can be cleaned up. - The path parsing doesn't handle comments or extra whitespace in the robots.txt file, which might lead to invalid paths being stored.
- Your success message doesn't show the actual URL or status code, making it hard to verify results.
Here's the fixed and improved code:
import urllib.request import urllib.error # Uncomment this line to let users input their own URL # url = input("Input Url:\n") url = 'https://stackoverflow.com/robots.txt' raw_robots = urllib.request.urlopen(url) robots = raw_robots.read().decode('utf-8') result_data_set = {"Disallowed": [], "Allowed": []} for line in robots.split("\n"): # Clean up the line and skip empty lines/comments cleaned_line = line.strip() if not cleaned_line or cleaned_line.startswith('#'): continue if cleaned_line.startswith('Allow:'): # Extract path, ignoring comments after the path path = cleaned_line.split(': ', 1)[1].split('#')[0].strip() result_data_set["Allowed"].append(path) elif cleaned_line.startswith('Disallow:'): path = cleaned_line.split(': ', 1)[1].split('#')[0].strip() result_data_set["Disallowed"].append(path) print("Parsed robots.txt rules:\n", result_data_set) base_url = 'https://stackoverflow.com' # Check Allowed paths print("\n--- Checking Allowed Paths ---") for path in result_data_set["Allowed"]: full_url = base_url + path try: response = urllib.request.urlopen(full_url) print(f"✅ {full_url} - Status Code: {response.getcode()}") except urllib.error.HTTPError as e: print(f"❌ {full_url} - HTTP Error: {e.code}") except urllib.error.URLError as e: print(f"⚠️ {full_url} - Connection Error: {e.reason}") # Check Disallowed paths print("\n--- Checking Disallowed Paths ---") for path in result_data_set["Disallowed"]: full_url = base_url + path try: response = urllib.request.urlopen(full_url) print(f"✅ {full_url} - Status Code: {response.getcode()}") except urllib.error.HTTPError as e: print(f"❌ {full_url} - HTTP Error: {e.code}") except urllib.error.URLError as e: print(f"⚠️ {full_url} - Connection Error: {e.reason}")
Key improvements made:
- Cleaner imports: Removed unused and duplicate imports to keep things lean.
- Robust line processing: Skips empty lines and comments, and properly extracts paths even if they have trailing comments (like
Allow: /abc # some note). - Correct path iteration: Now loops directly over the paths in each array (
AllowedandDisallowed) instead of dictionary keys. - Clearer output: Shows the full URL, status code, and uses emojis to make results easy to scan at a glance.
- Separated checks: Groups results for allowed and disallowed paths to keep output organized.
A quick note: Some robots.txt files use wildcards (like /abc*) or relative paths that your current code won't handle, but this version fixes the core functionality you asked for.
内容的提问来源于stack exchange,提问作者Blank95
相关产品推荐
相关产品推荐

