You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取免费代理列表网站无法获取代理,求无额外库的Python解决方案

Fixing Proxy Scraping with Regex (No External Libraries Needed)

I see the issue with your code—your regex uses capturing groups, which causes re.findall() to return tuples of matched segments instead of full, usable proxy strings. Let’s break down the problem and fix it:

The Root Cause

Your current regex r'([0-9]{1,3}\.){3}[0-9]{1,3}(:[0-9]{2,4})?' includes two capturing groups (the parentheses around parts of the pattern). When re.findall() encounters capturing groups, it returns a list of tuples containing each group’s match, not the complete combined proxy address. That’s why you’re seeing output like [('192.168.', ':8080'), ...] instead of full IP:Port strings.

The Fix: Non-Capturing Groups

Change the capturing groups to non-capturing by adding ?: inside the parentheses. This tells regex to use the groups for matching logic without storing them as separate values. Here’s the corrected regex:

r'(?:[0-9]{1,3}\.){3}[0-9]{1,3}(?::[0-9]{2,4})?'

Corrected Full Code

I also fixed a small typo in your User-Agent (Cafari → Safari) to avoid potential blocking, and added a basic error handler:

import requests
import re

url = 'https://free-proxy-list.net/'
headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_5) AppleWebKit/537.36 (KHTML, like Gecko) Safari/537.36'}

try:
    source = requests.get(url, headers=headers, timeout=10).text
    # Use non-capturing groups to get full proxy strings
    proxies = re.findall(r'(?:[0-9]{1,3}\.){3}[0-9]{1,3}(?::[0-9]{2,4})?', source)
    # Filter out any empty strings from partial matches
    proxies = [proxy for proxy in proxies if proxy]
    print(proxies)
except Exception as e:
    print(f"Error encountered: {e}")

Quick Extra Tip

If you want to target only proxies from the main table (and ignore random IPs elsewhere on the page), you can adjust the regex to match within table cells:

# Matches IP and port pairs directly from the table rows
proxy_matches = re.findall(r'<td>(?:[0-9]{1,3}\.){3}[0-9]{1,3}</td><td>([0-9]{2,4})</td>', source)
proxies = [f"{ip}:{port}" for ip, port in proxy_matches]

Note that this is slightly more fragile if the website’s HTML structure changes, but it’s a way to narrow down results without external libraries.

内容的提问来源于stack exchange,提问作者wished

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:15:54