You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python筛选并下载SEC EDGAR中的10-K文件

筛选SEC 10-K文件并批量下载

问题说明

已爬取指定股票(如AMZN)的SEC filings数据,但筛选type为10-K的链接时失败,遍历后仍获取所有类型链接。同时需要将筛选出的10-K文件批量下载到指定文件夹,不希望按年份拆分存储。

原代码如下:

from urllib.request import urlopen
import certifi
import json

response = urlopen("https://financialmodelingprep.com/api/v3/sec_filings/AMZN?page=0&apikey=aa478b6f376879bc58349bd2a6f9d5eb", cafile=certifi.where())
data = response.read().decode("utf-8")
print (json.loads(data))
list = []

for p_id in data:
    if p_id['type'] == '10-K':
        list.append((p_id['finalLink']))

print(list)

运行后仍输出所有类型的文件信息,无法筛选出目标10-K链接。

错误原因

核心问题是遍历了字符串而非解析后的JSON列表:

  • data是解码后的字符串,遍历data会逐个遍历每个字符,而非遍历 filings 条目
  • 解析后的JSON数据(json.loads(data))未被赋值给变量,后续遍历完全未用到有效数据

修正后的筛选代码

from urllib.request import urlopen
import certifi
import json

# 请求并解析数据
response = urlopen("https://financialmodelingprep.com/api/v3/sec_filings/AMZN?page=0&apikey=aa478b6f376879bc58349bd2a6f9d5eb", cafile=certifi.where())
data_str = response.read().decode("utf-8")
filings = json.loads(data_str)  # 解析为JSON列表

# 筛选10-K链接
ten_k_links = []
for filing in filings:
    if filing.get('type') == '10-K':  # 用get避免key不存在报错
        ten_k_links.append(filing['finalLink'])

print("筛选出的10-K链接:")
for link in ten_k_links:
    print(link)

批量下载到指定文件夹代码

添加文件下载逻辑,直接保存到目标文件夹,不按年份拆分:

from urllib.request import urlopen, Request
import certifi
import json
import os

# 配置目标文件夹
target_folder = "sec_10k_files"
os.makedirs(target_folder, exist_ok=True)  # 文件夹不存在则创建

# 请求并解析数据
response = urlopen("https://financialmodelingprep.com/api/v3/sec_filings/AMZN?page=0&apikey=aa478b6f376879bc58349bd2a6f9d5eb", cafile=certifi.where())
data_str = response.read().decode("utf-8")
filings = json.loads(data_str)

# 筛选并下载
for filing in filings:
    if filing.get('type') == '10-K':
        link = filing['finalLink']
        # 从链接提取文件名
        filename = os.path.basename(link)
        save_path = os.path.join(target_folder, filename)
        
        # 发起请求下载(添加User-Agent避免被SEC拦截)
        req = Request(link, headers={'User-Agent': 'Mozilla/5.0'})
        with urlopen(req, cafile=certifi.where()) as file_response, open(save_path, 'wb') as f:
            f.write(file_response.read())
        print(f"已下载:{save_path}")

注意事项

  • 添加User-Agent是因为SEC服务器会拦截无标识的请求,避免被拒绝访问
  • os.makedirs(target_folder, exist_ok=True)确保目标文件夹存在,不会因文件夹已存在报错
  • 用filing.get('type')代替直接索引,避免部分数据缺少type字段导致报错

内容的提问来源于stack exchange,提问作者joshijos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 16:30:47