如何处理Proxybroker生成的Selenium兼容代理文本文件
处理Proxybroker代理格式适配Selenium的方法
以下两种方法可快速将Proxybroker输出的.txt文件转换成Selenium可用格式,同时保留HTTP/HTTPS协议信息:
方法一:Python脚本处理
适合有Python环境的用户,逻辑清晰且易调整:
# 读取原始代理文件,处理后写入新文件 input_file = "proxies.txt" output_file = "formatted_proxies.txt" with open(input_file, 'r', encoding='utf-8') as f_in, open(output_file, 'w', encoding='utf-8') as f_out: for line in f_in: line = line.strip() if not line: continue # 去除首尾的<和> cleaned_line = line.strip('< >') # 分割协议信息与代理地址 parts = cleaned_line.split('] ') if len(parts) != 2: continue protocol_part, proxy_addr = parts # 提取协议并清理冗余描述(如: High) protocols = protocol_part.split('[')[1].split(', ') cleaned_protocols = [p.split(':')[0].strip() for p in protocols] # 为每个协议生成对应代理格式 for proto in cleaned_protocols: f_out.write(f"{proto}://{proxy_addr}\n")
运行脚本后,formatted_proxies.txt的内容示例:
HTTP://148.76.97.250:80 HTTP://47.88.62.42:80 HTTP://107.173.153.197:7777 HTTPS://107.173.153.197:7777 ...
该格式可直接用于Selenium代理配置,示例:
from selenium import webdriver from selenium.webdriver.common.proxy import Proxy, ProxyType proxy = Proxy() proxy.proxy_type = ProxyType.MANUAL proxy.http_proxy = "HTTP://148.76.97.250:80" proxy.ssl_proxy = "HTTPS://107.173.153.197:7777" options = webdriver.ChromeOptions() options.proxy = proxy driver = webdriver.Chrome(options=options)
方法二:命令行sed工具处理(Linux/macOS)
适合熟悉命令行的用户,无需编写脚本:
执行以下命令,将proxies.txt处理后输出到formatted_proxies.txt:
sed -E 's/^<Proxy [^[]+\[([^]]+)\] ([0-9.:]+)>$/\1,\2/' proxies.txt | sed -E 's/: High//g' | sed -E 's/, /\n/g' | awk -F',' '{for(i=1;i<NF;i++) print $i"://"$NF}' > formatted_proxies.txt
命令说明:
- 提取协议部分与代理地址,用逗号分隔
- 去除协议中的": High"冗余描述
- 将多协议拆分为单独行
- 为每个协议拼接代理地址,生成最终格式
处理后的文件内容与Python脚本输出一致,可直接用于Selenium。
内容的提问来源于stack exchange,提问作者chrumont velistitonk
相关产品推荐
相关产品推荐

