使用Python正则解析Requests响应时返回None的问题求助
Hey there! Let's dig into why your regex isn't matching the mining domain list you're fetching with Requests, and fix this step by step.
First off, the URL you're hitting returns a JSON object, not raw HTML or text. The actual domain list is tucked inside the content field of that JSON, and your initial approach missed parsing this structure entirely. Plus, your regex was targeting the wrong HTML pattern—those \u003e/\u003c escape sequences don't appear in the actual content once you parse the JSON correctly.
1. Parse the JSON Response Properly
When you use Requests to fetch the URL, you need to first decode the JSON to get access to the file content:
import requests import re url = 'https://gitlab.com/ZeroDot1/CoinBlockerLists/blob/master/list_browser.txt?format=json&viewer=simple' response = requests.get(url) response_data = response.json() # Decode the JSON response file_content_html = response_data['content'] # Extract the HTML containing the domains
2. Extract the Domain Text from <code> Tags
All the mining domains live inside a <code> block in the HTML content. Let's pull that out first:
# Match everything inside <code>...</code>, including newlines code_block_match = re.search(r'<code>(.*?)</code>', file_content_html, re.DOTALL) if not code_block_match: print("Couldn't find the code block with domains!") exit() domain_raw_text = code_block_match.group(1)
The re.DOTALL flag makes . match newline characters, which is critical since the domain list is multi-line.
3. Extract Individual Domains
Now you have a string where each line is a single domain. You can either split by lines (simplest) or use a regex to validate each entry:
# Method 1: Split by lines, filter out empty lines mining_domains = [line.strip() for line in domain_raw_text.split('\n') if line.strip()] # Method 2: Regex to match valid domain-like entries (optional, for stricter filtering) domain_pattern = re.compile(r'^\S+$', re.MULTILINE) mining_domains = domain_pattern.findall(domain_raw_text)
Full Working Code
Putting it all together, here's a robust script that fetches and extracts the domains:
import requests import re def fetch_mining_domains(): url = 'https://gitlab.com/ZeroDot1/CoinBlockerLists/blob/master/list_browser.txt?format=json&viewer=simple' try: response = requests.get(url) response.raise_for_status() # Catch HTTP errors (404, 500, etc.) response_data = response.json() file_content_html = response_data['content'] # Extract the code block with domains code_block_match = re.search(r'<code>(.*?)</code>', file_content_html, re.DOTALL) if not code_block_match: return [] domain_raw_text = code_block_match.group(1) # Get all non-empty lines as domains return [line.strip() for line in domain_raw_text.split('\n') if line.strip()] except Exception as e: print(f"Error fetching domains: {str(e)}") return [] # Test it out domains = fetch_mining_domains() print(f"Found {len(domains)} mining domains. Here are the first 5:") for domain in domains[:5]: print(f"- {domain}")
- No JSON Parsing: You tried running regex on the raw JSON string instead of extracting the
contentfield first. The outer JSON structure was blocking your regex from finding the right patterns. - Incorrect Regex Target: Your regex looked for
\u003e...\u003c/, but once you parse the JSON, those escape sequences are converted to actual</>characters. Worse, the domains aren't wrapped in that pattern anyway—they're just plain lines in a code block. - Unnecessary Byte/File Handling: You tried matching raw bytes or writing to a file, but the issue was never with encoding—it was with how you accessed and processed the content.
Now you can take the mining_domains list and add those entries to your firewall blacklist!
内容的提问来源于stack exchange,提问作者ThatOneGuy

