You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium或requests下载NYSCEF平台的受保护PDF?

无法通过Python下载纽约州法院NYSCEF网站的受保护PDF

目标文档URL:

https://iapps.courts.state.ny.us/nyscef/ViewDocument?docIndex=cdHe_PLUS_DaUdFKcTLzBtSo6zw==

使用Python的requests.get()或Selenium访问该页面时,遇到以下问题:

  • 用requests请求返回403 Forbidden响应
  • 用Selenium打开页面后显示空白,找不到<embed>标签

已尝试的方法

使用requests的代码

import requests

url = "https://iapps.courts.state.ny.us/nyscef/ViewDocument?docIndex=..."
headers = {
    "User-Agent": "Mozilla/5.0",
    "Referer": "https://iapps.courts.state.ny.us/nyscef/"
}
response = requests.get(url, headers=headers)
print(response.status_code)  # 始终返回403

使用SeleniumBase的代码

from seleniumbase import SB

with SB(headless=False) as sb:
    sb.open(url)
    sb.wait(5)
    try:
        embed = sb.find_element("embed")
        print(embed.get_attribute("src"))
    except Exception as e:
        print("❌ 未找到embed标签", e)

以上方法均无效。

完整参考代码

from seleniumbase import SB
import requests
import os
import time

def download_pdf_with_selenium_and_requests():
    # 目标文档URL
    doc_url = "https://iapps.courts.state.ny.us/nyscef/ViewDocument?docIndex=cdHe_PLUS_DaUdFKcTLzBtSo6zw=="

    # 设置下载目录
    download_dir = os.path.join(os.getcwd(), "downloads")
    os.makedirs(download_dir, exist_ok=True)
    filename = os.path.join(download_dir, "NYSCEF_Document.pdf")

    with SB(headless=True) as sb:
        # 步骤1:导航到文档页面(使用浏览器会话)
        sb.open(doc_url)
        time.sleep(5)  # 等待重定向/ cookie设置完成

        # 步骤2:获取实际PDF的<embed src>
        try:
            embed = sb.find_element("embed")
            pdf_url = embed.get_attribute("src")
            print(f"找到PDF URL: {pdf_url}")
        except Exception as e:
            print(f"未找到<embed>标签: {e}")
            return

        # 步骤3:从Selenium会话中提取cookie
        selenium_cookies = sb.driver.get_cookies()
        session = requests.Session()
        for cookie in selenium_cookies:
            session.cookies.set(cookie['name'], cookie['value'])

        # 步骤4:使用带cookie的requests下载PDF
        headers = {
            "User-Agent": "Mozilla/5.0",
            "Referer": doc_url
        }

        response = session.get(pdf_url, headers=headers)
        if response.status_code == 200 and "application/pdf" in response.headers.get("Content-Type", ""):
            with open(filename, "wb") as f:
                f.write(response.content)
            print(f"PDF已保存为: {filename}")
        else:
            print(f"PDF下载失败。状态码: {response.status_code}")
            print(f"内容类型: {response.headers.get('Content-Type')}")
            print(f"最终URL: {response.url}")

if __name__ == "__main__":
    download_pdf_with_selenium_and_requests()

执行结果

No <embed> tag found: Message: 
 Element {embed} was not present after 10 seconds!

内容的提问来源于stack exchange,提问作者Daremitsu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 14:25:11