You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Undetected Selenium Chrome Python脚本高数据用量优化求助

优化Selenium Chrome脚本流量消耗的方案

针对你10分钟消耗2.5GB流量的问题,结合你的脚本逻辑,从参数优化、执行逻辑、资源获取方式三个方向给你具体的优化方案:

一、补充ChromeOptions流量优化参数

你已经配置了部分禁用项,还可以添加以下参数进一步减少不必要的网络请求:

options.add_argument("--disable-background-downloads")  # 禁用后台下载任务
options.add_argument("--disable-default-apps")  # 禁用默认应用加载
options.add_argument("--disable-notifications")  # 阻止通知权限请求
options.add_argument("--safebrowsing-disable-auto-update")  # 禁用安全浏览库自动更新
options.add_argument("--no-default-browser-check")  # 跳过默认浏览器检测
options.add_argument("--no-first-run")  # 跳过首次运行初始化流程
options.add_argument("--disable-search-engine-choice-screen")  # 跳过搜索引擎选择界面
options.add_argument("--blink-settings=imagesEnabled=false")  # 更彻底禁用图片加载(部分场景比--disable-images效果更好)
options.add_argument("--disk-cache-size=104857600")  # 设置100MB磁盘缓存,避免重复下载静态资源
options.add_argument("--force-cache")  # 强制使用本地缓存
options.add_argument("--disable-features=LazyFrameLoading,LazyImageLoading")  # 禁用懒加载特性,减少意外触发的请求

注意:如果你的PDF下载依赖页面JS触发,需要注释掉--disable-javascript参数。

二、修改执行逻辑:复用浏览器实例

你当前每次循环都关闭驱动并杀死Chrome进程,重新启动会重复加载Chrome基础组件、配置文件等,这是巨大的流量浪费。改成复用一个浏览器实例,用标签页处理不同URL:

from selenium import webdriver
from selenium.common.exceptions import NoSuchElementException, NoSuchFrameException, TimeoutException, ElementNotInteractableException
import undetected_chromedriver as uc
import time

# 初始化一次浏览器
options = webdriver.ChromeOptions()
# 原有参数
options.add_argument("--disable-popup-blocking")
options.add_argument("--disable-images")
options.add_argument("--disable-extensions")
options.add_argument("--disable-component-extensions-with-background-pages")
options.add_argument("--disable-features=Translate")
options.add_argument("--mute-audio")
options.add_argument("--disable-background-networking")
options.add_argument("--disable-extensions--disable-sync")
options.add_argument("--disable-component-update")
options.add_argument("--disable-features=PrivacySandboxSettings4")
options.add_argument("--disable-javascript")
# 补充的优化参数
options.add_argument("--disable-background-downloads")
options.add_argument("--disable-default-apps")
options.add_argument("--disable-notifications")
options.add_argument("--safebrowsing-disable-auto-update")
options.add_argument("--no-default-browser-check")
options.add_argument("--no-first-run")
options.add_argument("--disable-search-engine-choice-screen")
options.add_argument("--blink-settings=imagesEnabled=false")
options.add_argument("--disk-cache-size=104857600")
options.add_argument("--force-cache")
options.add_argument("--disable-features=LazyFrameLoading,LazyImageLoading")

driver = uc.Chrome(options=options)
driver.maximize_window()

# 替换成你的实际URL列表
urls = ["https://www.website.com/"] * 50  

for url in urls:
    try:
        # 打开新标签页
        driver.execute_script("window.open('');")
        driver.switch_to.window(driver.window_handles[-1])
        driver.get(url)
        
        # 执行你的PDF下载逻辑(替换成你原来的# do something代码)
        time.sleep(3)
        
        # 关闭当前标签页并切回主标签页
        driver.close()
        driver.switch_to.window(driver.window_handles[0])
    except Exception as e:
        print(f"处理URL {url} 出错: {e}")
        # 出错时清理异常标签页
        if len(driver.window_handles) > 1:
            driver.close()
            driver.switch_to.window(driver.window_handles[0])

# 所有任务完成后统一关闭浏览器
driver.quit()

这种方式避免了重复初始化浏览器的大量冗余请求,能大幅降低流量消耗。

三、终极优化:直接请求PDF资源

如果能从页面中提取到PDF的直接下载链接,完全可以不用Selenium,改用requests库直接下载——不需要加载页面的HTML、CSS、字体等无关资源,流量消耗会降到最低:

import requests
import os

# 替换成实际的PDF直链列表
pdf_urls = ["https://example.com/doc1.pdf", "https://example.com/doc2.pdf"]  
save_dir = "./downloaded_pdfs/"
os.makedirs(save_dir, exist_ok=True)

for idx, pdf_url in enumerate(pdf_urls):
    try:
        # 流式下载避免占用过多内存
        response = requests.get(pdf_url, stream=True)
        response.raise_for_status()
        with open(f"{save_dir}document_{idx+1}.pdf", "wb") as f:
            for chunk in response.iter_content(chunk_size=8192):
                f.write(chunk)
        print(f"已下载 {pdf_url}")
    except Exception as e:
        print(f"下载 {pdf_url} 失败: {e}")

如果需要登录验证,可以把Selenium获取到的Cookie传给requests,实现免交互下载,这是流量优化的最优解。

内容的提问来源于stack exchange,提问作者Knockoutpie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 13:10:02