Selenium脚本入口点运行报[Errno 61]连接拒绝但功能正常
问题
我编写了一段Python脚本,用于抓取Google Flights的航班数据并写入文本文档。在终端用python3 code.py运行时完全正常,无任何报错;但将其添加到入口点(如cron、systemd等后台启动方式)运行时,会抛出以下错误:
Max retries exceeded with url: /session/2070f9827b90fc53d25392991e7b1855/url (Caused by NewConnectionError('<urllib3.connection.HTTPConnection object at 0x7fef3d47d520>: Failed to establish a new connection: [Errno 61] Connection refused'))
报错出现在browser.get(url)这一行,但奇怪的是,即使抛出该错误,脚本仍能正常工作,数据可以正确写入目标文件。完整代码如下:
from selenium import webdriver from time import time, sleep from datetime import datetime from selenium.webdriver.chrome.service import Service from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.relative_locator import locate_with from selenium.webdriver.common.keys import Keys from selenium.webdriver.chrome.options import Options import os #config headless mode options = Options() options.headless = True options.add_argument('--window-size=1920,1200') #Give path to chrome driver using service argument so it doesn't throw the path deprecation warning script_dir = os.path.dirname(__file__) #<-- absolute dir the script is in chromedriver_path = '/path/to/chromedriver' abs_chromedriver_path = os.path.join(script_dir, chromedriver_path) driver_service = Service(executable_path = abs_chromedriver_path) browser = webdriver.Chrome(options=options,service = driver_service) #all variables url = 'https://www.google.com/travel/flights' clickList = [ "//input[@aria-label='Departure']", "//div[@data-iso='2022-09-06']//div[@role='button']", "//div[@data-iso='2022-10-16']//div[@role='button']", "//div[@class='WXaAwc']//button" ] #Click function def browserClick(xPath): browser.find_element(By.XPATH, xPath).click() #Enter function def browserEnter(xPath): WebDriverWait(browser, 20).until(EC.element_to_be_clickable((By.XPATH, xPath))).send_keys(Keys.ENTER) #Type function def browserType(xPath,phrase): browserClick(xPath) WebDriverWait(browser, 20).until(EC.element_to_be_clickable((By.XPATH, xPath))).send_keys(phrase) browserEnter(xPath) #Average function def getAverage(lst): return sum(lst) / len(lst) def run(): browser.get(url) browserType("//div[@aria-label='Enter your origin']//preceding-sibling::div[1]//input", "Barcelona") browserType("//div[@aria-label='Enter your destination']//preceding-sibling::div[1]//input", "JFK") for click in clickList: browserClick(click) sleep(0.5) sleep(10) mainList = browser.find_elements(By.XPATH, '//ul[@class="Rk10dc"]//li') mainList.pop() prices = [] i = 1 #Time stamp now = datetime.now() timeTit = str(now) dt_string = now.strftime('%d/%m/%Y %H:%M') priceTime = now.strftime('%H:%M') #Setup backup file #script_dir = os.path.dirname(__file__) #<-- absolute dir the script is in rel_path_backup = 'backups/BackupFile.txt' abs_backup_path = os.path.join(script_dir, rel_path_backup) backupFile = open(abs_backup_path,'a') backupFile.write('\nBelow data scraped at: ' + dt_string + '\n') #Scrape flight data, distribute to python lists for element in mainList: #write whole item to backup data item = element.text backupFile.write('\n' + item + '\n') #parse item for price tempList = item.splitlines() #Check for stops if 'Nonstop' in tempList: priceSymInt = tempList[9] else: priceSymInt = tempList[10] priceInt = priceSymInt.replace('$','') if ',' in priceInt: priceInt = priceInt.replace(',','') prices.append(int(priceInt)) i += 1 #Close after all items have been processed backupFile.close() #Calculate average price avgPrice = getAverage(prices) rel_path_prices = 'prices/PricesFile.txt' abs_prices_path = os.path.join(script_dir, rel_path_prices) priceFile = open(abs_prices_path,'a') priceFile.write(priceTime + ',' + str(avgPrice) + '\n') priceFile.close() browser.close() browser.quit() run()
原因分析
这个错误本质是Selenium与ChromeDriver之间的HTTP通信连接失败,终端运行正常但入口点运行出错的核心差异在于环境:
- 后台环境权限/资源限制:终端运行时有完整的用户环境变量和网络权限,而入口点启动的进程可能缺少必要权限,或系统资源不足,导致ChromeDriver无法及时与浏览器实例建立连接。
- Headless模式隐性问题:后台环境中,仅开启
headless和窗口尺寸参数不足以让Chrome稳定运行,缺少沙箱禁用、GPU禁用等必要参数时,浏览器启动后可能无法正常响应Selenium指令,触发连接拒绝,但后续浏览器实例恢复正常,所以脚本仍能继续执行。 - 启动超时差异:终端环境下Chrome启动速度快,Selenium能及时建立连接;后台环境启动延迟高,Selenium超时后抛出错误,但之后浏览器实例完成启动,不影响后续流程。
解决办法
1. 补充Headless模式必备参数
给ChromeOptions添加后台运行必需的参数,消除环境差异:
options = Options() options.headless = True options.add_argument('--window-size=1920,1200') # 添加以下参数 options.add_argument('--no-sandbox') # 关闭沙箱,后台运行Chrome必备 options.add_argument('--disable-dev-shm-usage') # 避免/dev/shm临时空间不足 options.add_argument('--disable-gpu') # 无头模式无需GPU加速 options.add_argument('--remote-debugging-port=9222') # 指定固定调试端口,避免端口冲突 options.add_argument('--disable-extensions') # 禁用扩展,减少启动负载
2. 确保Chrome与ChromeDriver版本完全匹配
终端运行正常不代表后台环境的Chrome版本与你的ChromeDriver一致,执行google-chrome --version查看后台Chrome版本,下载对应版本的ChromeDriver替换现有文件。
3. 延长Selenium超时时间
给Chrome足够的启动时间,避免因后台启动慢导致的连接超时:
driver_service = Service(executable_path=abs_chromedriver_path) # 延长服务启动超时 driver_service.start() browser = webdriver.Chrome(options=options, service=driver_service) browser.set_page_load_timeout(30) # 延长页面加载超时至30秒
4. 统一入口点的环境变量与权限
- 确保入口点运行的用户与终端用户一致,或拥有访问ChromeDriver、
backups/、prices/目录的权限。 - 在入口点脚本中显式设置
PATH环境变量,确保能找到Chrome和ChromeDriver:# 比如在cron任务中添加 PATH=/usr/local/bin:/usr/bin:/bin
5. 捕获特定错误(临时兼容方案)
如果不想调整环境,可在browser.get(url)处捕获连接拒绝错误,不影响后续流程:
def run(): try: browser.get(url) except Exception as e: # 仅忽略连接拒绝类错误,其他错误正常抛出 if "Connection refused" in str(e): pass else: raise e # 后续代码不变
内容的提问来源于stack exchange,提问作者Fox

