Undetected Selenium Chrome Python脚本高数据用量优化求助
优化Selenium Chrome脚本流量消耗的方案
针对你10分钟消耗2.5GB流量的问题,结合你的脚本逻辑,从参数优化、执行逻辑、资源获取方式三个方向给你具体的优化方案:
一、补充ChromeOptions流量优化参数
你已经配置了部分禁用项,还可以添加以下参数进一步减少不必要的网络请求:
options.add_argument("--disable-background-downloads") # 禁用后台下载任务 options.add_argument("--disable-default-apps") # 禁用默认应用加载 options.add_argument("--disable-notifications") # 阻止通知权限请求 options.add_argument("--safebrowsing-disable-auto-update") # 禁用安全浏览库自动更新 options.add_argument("--no-default-browser-check") # 跳过默认浏览器检测 options.add_argument("--no-first-run") # 跳过首次运行初始化流程 options.add_argument("--disable-search-engine-choice-screen") # 跳过搜索引擎选择界面 options.add_argument("--blink-settings=imagesEnabled=false") # 更彻底禁用图片加载(部分场景比--disable-images效果更好) options.add_argument("--disk-cache-size=104857600") # 设置100MB磁盘缓存,避免重复下载静态资源 options.add_argument("--force-cache") # 强制使用本地缓存 options.add_argument("--disable-features=LazyFrameLoading,LazyImageLoading") # 禁用懒加载特性,减少意外触发的请求
注意:如果你的PDF下载依赖页面JS触发,需要注释掉--disable-javascript参数。
二、修改执行逻辑:复用浏览器实例
你当前每次循环都关闭驱动并杀死Chrome进程,重新启动会重复加载Chrome基础组件、配置文件等,这是巨大的流量浪费。改成复用一个浏览器实例,用标签页处理不同URL:
from selenium import webdriver from selenium.common.exceptions import NoSuchElementException, NoSuchFrameException, TimeoutException, ElementNotInteractableException import undetected_chromedriver as uc import time # 初始化一次浏览器 options = webdriver.ChromeOptions() # 原有参数 options.add_argument("--disable-popup-blocking") options.add_argument("--disable-images") options.add_argument("--disable-extensions") options.add_argument("--disable-component-extensions-with-background-pages") options.add_argument("--disable-features=Translate") options.add_argument("--mute-audio") options.add_argument("--disable-background-networking") options.add_argument("--disable-extensions--disable-sync") options.add_argument("--disable-component-update") options.add_argument("--disable-features=PrivacySandboxSettings4") options.add_argument("--disable-javascript") # 补充的优化参数 options.add_argument("--disable-background-downloads") options.add_argument("--disable-default-apps") options.add_argument("--disable-notifications") options.add_argument("--safebrowsing-disable-auto-update") options.add_argument("--no-default-browser-check") options.add_argument("--no-first-run") options.add_argument("--disable-search-engine-choice-screen") options.add_argument("--blink-settings=imagesEnabled=false") options.add_argument("--disk-cache-size=104857600") options.add_argument("--force-cache") options.add_argument("--disable-features=LazyFrameLoading,LazyImageLoading") driver = uc.Chrome(options=options) driver.maximize_window() # 替换成你的实际URL列表 urls = ["https://www.website.com/"] * 50 for url in urls: try: # 打开新标签页 driver.execute_script("window.open('');") driver.switch_to.window(driver.window_handles[-1]) driver.get(url) # 执行你的PDF下载逻辑(替换成你原来的# do something代码) time.sleep(3) # 关闭当前标签页并切回主标签页 driver.close() driver.switch_to.window(driver.window_handles[0]) except Exception as e: print(f"处理URL {url} 出错: {e}") # 出错时清理异常标签页 if len(driver.window_handles) > 1: driver.close() driver.switch_to.window(driver.window_handles[0]) # 所有任务完成后统一关闭浏览器 driver.quit()
这种方式避免了重复初始化浏览器的大量冗余请求,能大幅降低流量消耗。
三、终极优化:直接请求PDF资源
如果能从页面中提取到PDF的直接下载链接,完全可以不用Selenium,改用requests库直接下载——不需要加载页面的HTML、CSS、字体等无关资源,流量消耗会降到最低:
import requests import os # 替换成实际的PDF直链列表 pdf_urls = ["https://example.com/doc1.pdf", "https://example.com/doc2.pdf"] save_dir = "./downloaded_pdfs/" os.makedirs(save_dir, exist_ok=True) for idx, pdf_url in enumerate(pdf_urls): try: # 流式下载避免占用过多内存 response = requests.get(pdf_url, stream=True) response.raise_for_status() with open(f"{save_dir}document_{idx+1}.pdf", "wb") as f: for chunk in response.iter_content(chunk_size=8192): f.write(chunk) print(f"已下载 {pdf_url}") except Exception as e: print(f"下载 {pdf_url} 失败: {e}")
如果需要登录验证,可以把Selenium获取到的Cookie传给requests,实现免交互下载,这是流量优化的最优解。
内容的提问来源于stack exchange,提问作者Knockoutpie
相关产品推荐
相关产品推荐

