You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium下载PDF?手动过reCAPTCHA后下载无反应

解决reCAPTCHA验证后PDF未自动下载的问题

尝试从网址http://esaj.tjsp.jus.br/cjsg/getArquivo.do?conversationId=&cdAcordao=16548741下载PDF文件,设置time.sleep手动完成reCAPTCHA验证,但验证后下载并未启动。以下是使用的代码:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
import time

if __name__ == "__main__":

    url = "http://esaj.tjsp.jus.br/cjsg/getArquivo.do?conversationId=&cdAcordao=16548741"
    service = Service(executable_path='./chromedriver.exe')
    options = webdriver.ChromeOptions()
    
    options.add_experimental_option('prefs', {
         'download.default_directory': 'C:\\Users\\....',
         'download.prompt_for_download': False,
         "download.directory_upgrade": True,
         'plugins.always_open_pdf_externally': True})
    
    driver = webdriver.Chrome(service=service, options=options)
    try:
 
        driver.get(url)

        print('Acesso')

        time.sleep(180)
        
        print('Sucesso!')
        
        driver.implicitly_wait(10)
        
        print('Download concluido...')
        
        driver.quit()

    except Exception as e:
        print(f"Ocorreu um erro: {e}")
        driver.quit()

解决方法

  • 确认验证后的触发动作:reCAPTCHA验证通过后,部分网站需要手动点击确认/下载按钮才会触发下载。先手动操作一次观察页面是否有这类按钮,之后在代码中添加点击逻辑(需根据实际页面元素调整定位方式):

    from selenium.webdriver.common.by import By
    
    # 验证完成后,定位并点击下载触发按钮
    download_btn = driver.find_element(By.XPATH, "//button[contains(text(),'确认')]")
    download_btn.click()
    
  • 修正Chrome下载偏好配置:当前配置可能未完全禁用内置PDF查看器,调整配置如下:

    options.add_experimental_option('prefs', {
        'download.default_directory': 'C:\\Users\\你的实际路径',  # 替换为真实本地路径
        'download.prompt_for_download': False,
        'download.directory_upgrade': True,
        'plugins.plugins_list': [{"enabled": False, "name": "Chrome PDF Viewer"}],
        'plugins.always_open_pdf_externally': True
    })
    
  • 替换等待逻辑,准确判断下载状态:time.sleep和implicitly_wait无法精准判断下载是否完成,可通过检查目标目录的文件状态来确认:

    import os
    import glob
    
    def wait_for_download(download_dir, timeout=300):
        start_time = time.time()
        while time.time() - start_time < timeout:
            # 检查是否存在未完成的临时下载文件(.crdownload后缀)
            temp_files = glob.glob(os.path.join(download_dir, '*.crdownload'))
            if not temp_files:
                # 检查是否有已完成的PDF文件
                pdf_files = glob.glob(os.path.join(download_dir, '*.pdf'))
                if pdf_files:
                    return True
            time.sleep(1)
        return False
    
    # 在验证完成后调用该函数
    if wait_for_download('C:\\Users\\你的实际路径'):
        print('下载完成...')
    else:
        print('下载超时')
    
  • 处理网站的页面跳转/重新请求:部分网站验证后会刷新页面导致下载链接失效,可在验证完成后重新请求目标链接:

    time.sleep(5)  # 等待验证后的页面跳转
    driver.get(url)  # 重新发起下载请求
    

内容的提问来源于stack exchange,提问作者Claudio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 01:04:53