爬取免费代理列表网站无法获取代理,求无额外库的Python解决方案
I see the issue with your code—your regex uses capturing groups, which causes re.findall() to return tuples of matched segments instead of full, usable proxy strings. Let’s break down the problem and fix it:
The Root Cause
Your current regex r'([0-9]{1,3}\.){3}[0-9]{1,3}(:[0-9]{2,4})?' includes two capturing groups (the parentheses around parts of the pattern). When re.findall() encounters capturing groups, it returns a list of tuples containing each group’s match, not the complete combined proxy address. That’s why you’re seeing output like [('192.168.', ':8080'), ...] instead of full IP:Port strings.
The Fix: Non-Capturing Groups
Change the capturing groups to non-capturing by adding ?: inside the parentheses. This tells regex to use the groups for matching logic without storing them as separate values. Here’s the corrected regex:
r'(?:[0-9]{1,3}\.){3}[0-9]{1,3}(?::[0-9]{2,4})?'
Corrected Full Code
I also fixed a small typo in your User-Agent (Cafari → Safari) to avoid potential blocking, and added a basic error handler:
import requests import re url = 'https://free-proxy-list.net/' headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_5) AppleWebKit/537.36 (KHTML, like Gecko) Safari/537.36'} try: source = requests.get(url, headers=headers, timeout=10).text # Use non-capturing groups to get full proxy strings proxies = re.findall(r'(?:[0-9]{1,3}\.){3}[0-9]{1,3}(?::[0-9]{2,4})?', source) # Filter out any empty strings from partial matches proxies = [proxy for proxy in proxies if proxy] print(proxies) except Exception as e: print(f"Error encountered: {e}")
Quick Extra Tip
If you want to target only proxies from the main table (and ignore random IPs elsewhere on the page), you can adjust the regex to match within table cells:
# Matches IP and port pairs directly from the table rows proxy_matches = re.findall(r'<td>(?:[0-9]{1,3}\.){3}[0-9]{1,3}</td><td>([0-9]{2,4})</td>', source) proxies = [f"{ip}:{port}" for ip, port in proxy_matches]
Note that this is slightly more fragile if the website’s HTML structure changes, but it’s a way to narrow down results without external libraries.
内容的提问来源于stack exchange,提问作者wished

