You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取SEC Edgar模态框页脚内文档链接的问题

问题分析

你当前代码无法拿到对应链接的核心原因是:SEC Edgar的搜索结果页面为动态渲染页面,你使用requests.get()获取到的仅为未加载业务数据的静态HTML骨架,所有搜索结果、模态框内的链接元素都是后续浏览器执行JS代码调用后端接口拉取数据后才生成的,自然无法通过BeautifulSoup从静态源码中解析到。

你提供的对应DOM结构快照如下:
网站HTML代码中对应链接的快照

你当前的测试代码:

url = 'https://www.sec.gov/edgar/search/#/q=ex10&category=custom&forms=10-K%252C10-Q%252C8-K'
source_code = requests.get(url)
plain_text = source_code.text
soup = BeautifulSoup(plain_text)
    
for a in soup.find_all(id = "open-file"):
    print(a)
可直接落地的解决方案

这里提供两种简单的实现方案,按需选择即可:

  • 方案一:直接调用官方后端搜索接口(更推荐,无需渲染页面,速度快)
    SEC的搜索结果都是从公开的后端接口返回的,你可以直接请求接口拿到JSON格式的所有文档数据,不需要解析前端页面。示例代码如下:
    import requests
    import json
    
    headers = {
        "User-Agent": "填写你的标识,比如个人邮箱或者项目名,SEC要求请求带合理标识否则容易被拦截",
        "Content-Type": "application/json"
    }
    # 对应你的搜索条件的请求参数
    payload = {
        "q": "ex10",
        "category": "custom",
        "forms": ["10-K", "10-Q", "8-K"],
        "from": 0,
        "size": 100 # 单次返回结果数,可根据需求调整
    }
    response = requests.post("https://efts.sec.gov/LATEST/search-index", headers=headers, data=json.dumps(payload))
    data = response.json()
    
    # 遍历结果拼接htm链接
    for item in data["hits"]["hits"]:
        cik = item["_source"]["ciks"][0]
        accession_num = item["_source"]["adsh"].replace("-", "")
        file_name = item["_source"]["file"]
        # 拼出完整的htm文档链接
        htm_url = f"https://www.sec.gov/Archives/edgar/data/{cik}/{accession_num}/{file_name}"
        print(htm_url)
    
  • 方案二:使用模拟浏览器工具渲染页面后解析
    如果你需要和页面交互、复现前端点击模态框的逻辑,可以用Selenium等工具模拟浏览器运行,等页面完全渲染后再提取元素。示例代码如下:
    from selenium import webdriver
    from selenium.webdriver.common.by import By
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    
    driver = webdriver.Chrome() # 需要提前安装ChromeDriver和selenium库
    driver.get("https://www.sec.gov/edgar/search/#/q=ex10&category=custom&forms=10-K%252C10-Q%252C8-K")
    
    # 等待搜索结果加载完成
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "entity-name"))
    )
    
    # 示例:点击第一个结果触发模态框,再拿链接
    first_result = driver.find_element(By.CLASS_NAME, "preview-file")
    first_result.click()
    
    # 等待模态框的链接加载完成
    open_file_btn = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.ID, "open-file"))
    )
    print(open_file_btn.get_attribute("href"))
    
    driver.quit()
    

注意:SEC对爬虫请求频率有限制,建议添加至少1秒的请求间隔,同时请求头携带合理的User-Agent,避免被IP封禁。

内容的提问来源于stack exchange,提问作者Steve

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 16:27:03