Selenium配置ScraperAPI代理不生效,如何无需selenium-wire完成代理设置
问题核心原因
Chrome原生的--proxy-server启动参数不支持直接传入包含用户名、密码认证信息的代理地址,你之前取消注释后无法运行就是这个原因,不需要依赖selenium-wire也有两种成熟的解决方案。
方案1:IP白名单绑定(零代码修改,最简便)
- 登录你的ScraperAPI后台,找到IP白名单配置项,把你当前运行代码设备的公网出口IP添加到白名单
- 把代理地址修改为不带认证信息的格式即可直接运行:
PROXY = 'http://proxy-server.scraperapi.com:8001' options.add_argument('--proxy-server=%s' % PROXY)
这种方案不需要对原有代码做其他逻辑修改,适配性最高。
方案2:内置代理认证扩展(无需绑定IP,无额外依赖)
如果不想绑定固定IP,可以通过动态生成一个极小的Chrome代理认证扩展实现,启动浏览器时自动加载扩展完成认证,完整修改后的代码如下:
import time import os import sys import zipfile from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.chrome.service import Service from sys import platform from selenium.webdriver.support.ui import WebDriverWait from webdriver_manager.chrome import ChromeDriverManager from fake_useragent import UserAgent from dotenv import load_dotenv, find_dotenv WAIT = 10 load_dotenv(find_dotenv()) SCRAPER_API = os.environ.get("SCRAPER_API") PROXY_HOST = 'proxy-server.scraperapi.com' PROXY_PORT = 8001 PROXY_USER = 'scraperapi' PROXY_PASS = SCRAPER_API # 动态生成代理认证扩展 def create_proxy_auth_extension(proxy_host, proxy_port, proxy_username, proxy_password, scheme='http', plugin_path=None): if plugin_path is None: plugin_path = 'proxy_auth_plugin.zip' manifest_json = """ { "version": "1.0.0", "manifest_version": 2, "name": "Chrome Proxy", "permissions": [ "proxy", "tabs", "unlimitedStorage", "storage", "<all_urls>", "webRequest", "webRequestBlocking" ], "background": { "scripts": ["background.js"] }, "minimum_chrome_version":"22.0.0" } """ background_js = """ var config = { mode: "fixed_servers", rules: { singleProxy: { scheme: "%s", host: "%s", port: parseInt(%s) }, bypassList: ["localhost"] } }; chrome.proxy.settings.set({value: config, scope: "regular"}, function() {}); function callbackFn(details) { return { authCredentials: { username: "%s", password: "%s" } }; } chrome.webRequest.onAuthRequired.addListener( callbackFn, {urls: ["<all_urls>"]}, ['blocking'] ); """ % (scheme, proxy_host, proxy_port, proxy_username, proxy_password) with zipfile.ZipFile(plugin_path, 'w') as zp: zp.writestr("manifest.json", manifest_json) zp.writestr("background.js", background_js) return plugin_path # 生成扩展文件 proxy_auth_plugin = create_proxy_auth_extension(PROXY_HOST, PROXY_PORT, PROXY_USER, PROXY_PASS) srv=Service(ChromeDriverManager().install()) ua = UserAgent() userAgent = ua.random options = Options() # 注意老版headless模式不支持加载扩展,改用新版headless模式 options.add_argument('--headless=new') options.add_experimental_option ('excludeSwitches', ['enable-logging']) options.add_argument("start-maximized") options.add_argument('window-size=1920x1080') options.add_argument('--no-sandbox') options.add_argument('--disable-gpu') options.add_argument(f'user-agent={userAgent}') # 加载代理认证扩展 options.add_extension(proxy_auth_plugin) path = os.path.abspath (os.path.dirname (sys.argv[0])) if platform == "win32": cd = '/chromedriver.exe' elif platform == "linux": cd = '/chromedriver' elif platform == "darwin": cd = '/chromedriver' driver = webdriver.Chrome (service=srv, options=options) waitWebDriver = WebDriverWait (driver, 10) link = "https://whatismyipaddress.com/" driver.get (link) time.sleep(WAIT) soup = BeautifulSoup (driver.page_source, 'html.parser') tmpIP = soup.find("span", {"id": "ipv4"}) tmpP = soup.find_all("p", {"class": "information"}) for e in tmpP: tmpSPAN = e.find_all("span") for e2 in tmpSPAN: print(e2.text) print(tmpIP.text) driver.quit() # 运行结束后自动删除生成的扩展文件 os.remove(proxy_auth_plugin)
如果你的Chrome版本特别老不支持--headless=new,只能用旧headless模式的话,优先选择方案1的IP白名单配置即可。
内容的提问来源于stack exchange,提问作者Rapid1898
相关产品推荐
相关产品推荐

