Python 2.7中使用字典筛选状态码200的URL失败问题排查
Hey, let's break down why your code keeps outputting "po" instead of the URLs with 200 status codes you're looking for. There are a few critical issues in your implementation:
1. Type Mismatch Breaks the Comparison
You're converting the status code to a string when storing it as a dictionary key: req2 = str(req.status_code), but statuscode[0] is an integer (200). In Python, a string like "200" and the integer 200 are not equal—so your keys == statuscode[0] check will always return False, sending you straight to the else branch.
2. Dictionary Gets Overwritten, and Filter Logic Is in the Wrong Place
- Every time you make a successful request, you create a brand new dictionary with
dict = {req2 : req3}. This means you're throwing away all previous entries, so only the last successful request's data stays in the dictionary. - Your filter loop (
for keys, values in dict.iteritems():) sits outside the main subdomain loop, so it only runs once—after all subdomains are processed—on just that final dictionary entry.
3. Bonus: Potential Errors Waiting to Happen
- If all requests fail, the
dictvariable will never be defined, and your filter loop will throw aNameError. - Using
dictas a variable name is a bad idea—it overrides Python's built-indicttype, which can cause unexpected bugs later.
Fixed Code
Let's rewrite this to properly track all successful requests and filter the 200-status URLs correctly. Here's a cleaner approach using a list to store results:
import requests # Use a clear variable name instead of a list for a single value target_status = 200 # Store all successful (status_code, url) pairs successful_requests = [] for x in find_from_sublister(hostname): x2 = x.strip() url = "http://" + x2 try: # Add a timeout to avoid hanging on unresponsive subdomains req = requests.get(url, timeout=5) status_code = req.status_code print "\n{} {}".format(status_code, url) # Store the integer status code and URL successful_requests.append( (status_code, url) ) except requests.exceptions.RequestException as e: print "Can't make the request to this Subdomain {}: {}".format(url, str(e)) # Filter and print URLs with 200 status code print "\n--- URLs with Status Code 200 ---" found_200 = False for code, url in successful_requests: if code == target_status: print url found_200 = True if not found_200: print "No URLs returned a 200 status code."
If you prefer using a dictionary to group URLs by their status codes (which is useful if you want to check other codes later), here's that version:
import requests target_status = 200 # Key: integer status code, Value: list of URLs with that code status_to_urls = {} for x in find_from_sublister(hostname): x2 = x.strip() url = "http://" + x2 try: req = requests.get(url, timeout=5) status_code = req.status_code print "\n{} {}".format(status_code, url) # Add the URL to the corresponding status code list if status_code not in status_to_urls: status_to_urls[status_code] = [] status_to_urls[status_code].append(url) except requests.exceptions.RequestException as e: print "Can't make the request to this Subdomain {}: {}".format(url, str(e)) # Print 200-status URLs print "\n--- URLs with Status Code 200 ---" if target_status in status_to_urls: for url in status_to_urls[target_status]: print url else: print "No URLs returned a 200 status code."
These changes fix the type mismatch, preserve all successful requests, and place the filter logic where it needs to be to check all results.
内容的提问来源于stack exchange,提问作者Utkarsh Agrawal

