Docker中Python爬虫设置代理后NO_PROXY失效,Selenium无法正常工作
解决Docker中Scrapy+Selenium代理导致本地连接失败的问题
问题背景
在Docker容器中运行基于Scrapy和Selenium(Chrome)的爬虫时,未设置代理一切正常,但配置全局代理后Selenium无法正常工作。通过bash脚本设置的代理环境变量如下:
export HTTP_PROXY=http://user:password@p.webshare.io:80/ export HTTPS_PROXY=http://user:password@p.webshare.io:80/ export NO_PROXY=localhost,127.0.0.1 export no_proxy=localhost,127.0.0.1 python main.py
启用代理后,日志显示Selenium忽略了NO_PROXY配置,尝试通过代理连接本地ChromeDriver,导致404错误:
app_1 | 2023-01-16 21:59:47 [urllib3.connectionpool] DEBUG: http://localhost:58683 "POST /session/fb57da8f4b8b938c2ec744eb32cc2ca4/element HTTP/1.1" 404 883 app_1 | 2023-01-16 21:59:47 [selenium.webdriver.remote.remote_connection] DEBUG: Remote response: status=404 | data={"value":{"error":"no such element","message":"no such element: Unable to locate element: {\"method\":\"css selector\",\"selector\":\".mu-product-details-page\"}\n (Session info: headless chrome=108.0.5359.124)","stacktrace":"#0 0x5574a38762a3 \u003Cunknown\>\n#1 0x5574a3634f77 \u003Cunknown\>\n#2 0x5574a367180c \u003Cunknown\>\n#3 0x5574a3671a71 \u003Cunknown\>\n#4 0x5574a36ab734 \u003Cunknown\>\n#5 0x5574a3691b5d \u003Cunknown\>\n#6 0x5574a36a947c \u003Cunknown\>\n#7 0x5574a3691903 \u003Cunknown\>\n#8 0x5574a3664ece \u003Cunknown\>\n#9 0x5574a3665fde \u003Cunknown\>\n#10 0x5574a38c663e \u003Cunknown\>\n#11 0x5574a38c9b79 \u003Cunknown\>\n#12 0x5574a38ac89e \u003Cunknown\>\n#13 0x5574a38caa83 \u003Cunknown\>\n#14 0x5574a389f505 \u003Cunknown\>\n#15 0x5574a38ebca8 \u003Cunknown\>\n#16 0x5574a38ebe36 \u003Cunknown\>\n#17 0x5574a3907333 \u003Cunknown\>\n#18 0x7f52d0dc4609 start_thread\n"}} | headers=HTTPHeaderDict({'Content-Length': '883', 'Content-Type': 'application/json; charset=utf-8', 'cache-control': 'no-cache'})
尝试将带端口的localhost加入NO_PROXY,但端口随机变化,批量添加端口会导致列表过长无法生效。
解决方案:自定义RemoteConnection忽略代理
通过创建自定义RemoteConnection对象并设置ignore_proxy=True,可以让Selenium与本地ChromeDriver的通信绕过系统代理,仅让爬虫的网络请求走代理。
具体实现代码
方式1:直接初始化Chrome浏览器时传入自定义连接
from selenium import webdriver from selenium.webdriver.remote.remote_connection import RemoteConnection from selenium.webdriver.chrome.service import Service # 配置ChromeDriver的地址(默认是http://localhost:9515,根据实际情况调整) chrome_driver_url = "http://localhost:9515" # 创建自定义连接,忽略代理 custom_connection = RemoteConnection(chrome_driver_url, ignore_proxy=True) # 配置Chrome选项(根据你的需求添加参数,比如无头模式、禁用沙箱等) chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") # 配置ChromeDriver服务(如果需要指定驱动路径) service = Service(executable_path="/usr/local/bin/chromedriver") # 初始化Chrome浏览器,传入自定义连接 driver = webdriver.Chrome( service=service, options=chrome_options, remote_connection=custom_connection )
方式2:使用RemoteWebDriver连接本地驱动
如果你的代码是通过RemoteWebDriver连接本地ChromeDriver,写法如下:
from selenium import webdriver from selenium.webdriver.remote.remote_connection import RemoteConnection chrome_driver_url = "http://localhost:9515" custom_connection = RemoteConnection(chrome_driver_url, ignore_proxy=True) chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--headless=new") driver = webdriver.Remote( command_executor=custom_connection, options=chrome_options )
原理说明
RemoteConnection是Selenium与浏览器驱动通信的底层组件,设置ignore_proxy=True后,该组件会跳过系统代理配置,直接与本地ChromeDriver建立连接,避免了代理对本地通信的干扰。而Scrapy的网络请求仍然会遵循系统的代理配置,互不影响。
内容的提问来源于stack exchange,提问作者tjwnuk
相关产品推荐
相关产品推荐

