使用BeautifulSoup提取SEC Edgar模态框页脚内文档链接的问题
问题分析
你当前代码无法拿到对应链接的核心原因是:SEC Edgar的搜索结果页面为动态渲染页面,你使用requests.get()获取到的仅为未加载业务数据的静态HTML骨架,所有搜索结果、模态框内的链接元素都是后续浏览器执行JS代码调用后端接口拉取数据后才生成的,自然无法通过BeautifulSoup从静态源码中解析到。
你提供的对应DOM结构快照如下:
你当前的测试代码:
url = 'https://www.sec.gov/edgar/search/#/q=ex10&category=custom&forms=10-K%252C10-Q%252C8-K' source_code = requests.get(url) plain_text = source_code.text soup = BeautifulSoup(plain_text) for a in soup.find_all(id = "open-file"): print(a)
可直接落地的解决方案
这里提供两种简单的实现方案,按需选择即可:
- 方案一:直接调用官方后端搜索接口(更推荐,无需渲染页面,速度快)
SEC的搜索结果都是从公开的后端接口返回的,你可以直接请求接口拿到JSON格式的所有文档数据,不需要解析前端页面。示例代码如下:import requests import json headers = { "User-Agent": "填写你的标识,比如个人邮箱或者项目名,SEC要求请求带合理标识否则容易被拦截", "Content-Type": "application/json" } # 对应你的搜索条件的请求参数 payload = { "q": "ex10", "category": "custom", "forms": ["10-K", "10-Q", "8-K"], "from": 0, "size": 100 # 单次返回结果数,可根据需求调整 } response = requests.post("https://efts.sec.gov/LATEST/search-index", headers=headers, data=json.dumps(payload)) data = response.json() # 遍历结果拼接htm链接 for item in data["hits"]["hits"]: cik = item["_source"]["ciks"][0] accession_num = item["_source"]["adsh"].replace("-", "") file_name = item["_source"]["file"] # 拼出完整的htm文档链接 htm_url = f"https://www.sec.gov/Archives/edgar/data/{cik}/{accession_num}/{file_name}" print(htm_url) - 方案二:使用模拟浏览器工具渲染页面后解析
如果你需要和页面交互、复现前端点击模态框的逻辑,可以用Selenium等工具模拟浏览器运行,等页面完全渲染后再提取元素。示例代码如下:from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() # 需要提前安装ChromeDriver和selenium库 driver.get("https://www.sec.gov/edgar/search/#/q=ex10&category=custom&forms=10-K%252C10-Q%252C8-K") # 等待搜索结果加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "entity-name")) ) # 示例:点击第一个结果触发模态框,再拿链接 first_result = driver.find_element(By.CLASS_NAME, "preview-file") first_result.click() # 等待模态框的链接加载完成 open_file_btn = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "open-file")) ) print(open_file_btn.get_attribute("href")) driver.quit()
注意:SEC对爬虫请求频率有限制,建议添加至少1秒的请求间隔,同时请求头携带合理的User-Agent,避免被IP封禁。
内容的提问来源于stack exchange,提问作者Steve
相关产品推荐
相关产品推荐

