Python代码改造:将URL与拦截列表统计结果从列表存储改为字典存储
Convert Results to Dictionary Instead of List
Got it, let’s adjust your code to store the URL-to-red-count mapping in a dictionary—this will make it way easier to look up counts by specific URLs later. Here’s the tweaked version with clear explanations:
Key Adjustments Made:
- Swapped the result list for a dictionary where URLs are keys and their corresponding red counts are values.
- Simplified the
starts_with_hashfunction to be more concise while keeping the same logic. - Improved error handling to return the problematic URL alongside an error marker (instead of a generic 'booger' string), so you can track which requests failed.
- Added optional tracking for failed URLs if you need to debug issues down the line.
Modified Code
import requests import concurrent.futures import utils # Assuming this is your existing module def starts_with_hash(child: str) -> bool: # Keep lines that DON'T start with a # (simplified one-liner) return not child.startswith('#') def fetch_blocklist_count(session: requests.Session, url: str, timeout: int): try: with session.get(url, timeout=timeout) as response: # Filter out comment lines and count the remaining valid_lines = filter(starts_with_hash, response.text.splitlines()) red_count = len(list(valid_lines)) return url, red_count except requests.exceptions.RequestException as e: # Return the failed URL + error marker instead of a vague string return url, None # Use 'error' instead of None if you prefer a string def test_count(timeout: int=10): session = requests.Session() results_dict = {} # Empty dict to store URL:count pairs failed_urls = [] # Optional: track which URLs failed to load with concurrent.futures.ThreadPoolExecutor() as executor: futures = [] block_host = utils.build_source_list() for blocklist in block_host: if not blocklist.startswith('#'): futures.append(executor.submit(fetch_blocklist_count, session, blocklist, timeout)) # Process each completed task for future in concurrent.futures.as_completed(futures): url, result = future.result() if result is not None: # Success: add the URL and count to the dict results_dict[url] = result else: # Failure: log the URL for debugging failed_urls.append(url) # Return either just the dict, or both dict + failed list (your call) return results_dict, failed_urls # Quick test run (optional) if __name__ == "__main__": url_counts, failed = test_count() print("URL to Red Count Mapping:", url_counts) print("Failed to Fetch URLs:", failed)
How It Works:
- Dictionary Storage:
results_dictlets you directly look up counts withurl_counts["https://example.com/blocklist.txt"]instead of searching through a list. - Error Clarity: Returning
(url, None)on failure makes it obvious which URLs had issues, and thefailed_urlslist keeps them organized for debugging. - Cleaner Filter Logic: The simplified
starts_with_hashfunction does exactly what it says—we keep lines that aren’t comments.
If you don’t need to track failed URLs, you can remove that section entirely and just populate the dictionary with successful results.
内容的提问来源于stack exchange,提问作者uncrayon
相关产品推荐
相关产品推荐

