Ubuntu WSL环境下Airflow连接Selenium WebDriver失败的解决咨询
问题:Ubuntu WSL下Airflow使用Selenium爬取时无法连接chromedriver
在Ubuntu WSL环境运行Apache Airflow测试平台,使用Selenium进行网页爬取时出现以下错误:
File "/home/siva/.local/lib/python3.10/site-packages/selenium/webdriver/chromium/webdriver.py", line 89, in __init__ self.service.start() File "/home/siva/.local/lib/python3.10/site-packages/selenium/webdriver/common/service.py", line 105, in start raise WebDriverException("Can not connect to the Service %s" % self.path) selenium.common.exceptions.WebDriverException: Message: Can not connect to the Service /d/apache-airflow/dags/chromedriver.exe
解决方案
一、修复Selenium与chromedriver的连接问题
错误核心是你在WSL(Linux环境)中使用了Windows版本的chromedriver.exe,两者环境不兼容,按以下步骤处理:
安装Linux版Chrome浏览器和chromedriver
- 更新系统软件源:
sudo apt update - 安装Chromium浏览器(Ubuntu官方源中的Chrome兼容版):
sudo apt install chromium-browser - 安装对应版本的chromedriver:
sudo apt install chromium-chromedriver - 验证chromedriver路径:执行
which chromedriver,会输出类似/usr/bin/chromedriver的路径。
- 更新系统软件源:
修改Airflow DAG中的代码路径
将代码中指定的chromedriver路径替换为Linux版路径,Selenium 4+推荐使用Service类管理驱动:from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options # 配置无头模式(WSL无图形界面必须开启) chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--no-sandbox') chrome_options.add_argument('--disable-dev-shm-usage') # 初始化驱动 service = Service('/usr/bin/chromedriver') driver = webdriver.Chrome(service=service, options=chrome_options)确保执行权限
若出现权限错误,赋予chromedriver执行权限:sudo chmod +x /usr/bin/chromedriver
二、Airflow中其他可行的网页爬取方式
如果Selenium配置过于繁琐,可根据爬取场景选择以下更适配的方案:
1. Requests + BeautifulSoup(适合静态页面)
无需浏览器,直接发送HTTP请求并解析HTML,资源占用极低,适合爬取无动态渲染的页面:
import requests from bs4 import BeautifulSoup def scrape_static_page(): url = "https://example.com" response = requests.get(url) response.raise_for_status() # 捕获请求错误 soup = BeautifulSoup(response.text, 'html.parser') # 示例:提取页面标题 page_title = soup.title.string print(page_title)
2. Scrapy(适合大规模爬虫)
专为爬虫设计的框架,自带异步请求、数据管道、去重等功能,可通过Airflow的BashOperator调用Scrapy命令,或用PythonOperator直接集成爬虫逻辑:
# 示例:Airflow中调用Scrapy的BashOperator from airflow.operators.bash import BashOperator scrape_task = BashOperator( task_id='run_scrapy_spider', bash_command='cd /path/to/your/scrapy/project && scrapy crawl your_spider_name' )
3. Playwright(替代Selenium的现代方案)
支持多浏览器(Chrome/Firefox/Safari),安装和配置更简单,无头模式稳定性更高,在WSL中无需额外图形界面配置:
from playwright.sync_api import sync_playwright def scrape_with_playwright(): with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page() page.goto("https://example.com") # 示例:提取页面文本 page_text = page.text_content() print(page_text) browser.close()
内容的提问来源于stack exchange,提问作者Siva Rama krishna Kumar
相关产品推荐
相关产品推荐

