如何用Python筛选并下载SEC EDGAR中的10-K文件
筛选SEC 10-K文件并批量下载
问题说明
已爬取指定股票(如AMZN)的SEC filings数据,但筛选type为10-K的链接时失败,遍历后仍获取所有类型链接。同时需要将筛选出的10-K文件批量下载到指定文件夹,不希望按年份拆分存储。
原代码如下:
from urllib.request import urlopen import certifi import json response = urlopen("https://financialmodelingprep.com/api/v3/sec_filings/AMZN?page=0&apikey=aa478b6f376879bc58349bd2a6f9d5eb", cafile=certifi.where()) data = response.read().decode("utf-8") print (json.loads(data)) list = [] for p_id in data: if p_id['type'] == '10-K': list.append((p_id['finalLink'])) print(list)
运行后仍输出所有类型的文件信息,无法筛选出目标10-K链接。
错误原因
核心问题是遍历了字符串而非解析后的JSON列表:
data是解码后的字符串,遍历data会逐个遍历每个字符,而非遍历 filings 条目- 解析后的JSON数据(
json.loads(data))未被赋值给变量,后续遍历完全未用到有效数据
修正后的筛选代码
from urllib.request import urlopen import certifi import json # 请求并解析数据 response = urlopen("https://financialmodelingprep.com/api/v3/sec_filings/AMZN?page=0&apikey=aa478b6f376879bc58349bd2a6f9d5eb", cafile=certifi.where()) data_str = response.read().decode("utf-8") filings = json.loads(data_str) # 解析为JSON列表 # 筛选10-K链接 ten_k_links = [] for filing in filings: if filing.get('type') == '10-K': # 用get避免key不存在报错 ten_k_links.append(filing['finalLink']) print("筛选出的10-K链接:") for link in ten_k_links: print(link)
批量下载到指定文件夹代码
添加文件下载逻辑,直接保存到目标文件夹,不按年份拆分:
from urllib.request import urlopen, Request import certifi import json import os # 配置目标文件夹 target_folder = "sec_10k_files" os.makedirs(target_folder, exist_ok=True) # 文件夹不存在则创建 # 请求并解析数据 response = urlopen("https://financialmodelingprep.com/api/v3/sec_filings/AMZN?page=0&apikey=aa478b6f376879bc58349bd2a6f9d5eb", cafile=certifi.where()) data_str = response.read().decode("utf-8") filings = json.loads(data_str) # 筛选并下载 for filing in filings: if filing.get('type') == '10-K': link = filing['finalLink'] # 从链接提取文件名 filename = os.path.basename(link) save_path = os.path.join(target_folder, filename) # 发起请求下载(添加User-Agent避免被SEC拦截) req = Request(link, headers={'User-Agent': 'Mozilla/5.0'}) with urlopen(req, cafile=certifi.where()) as file_response, open(save_path, 'wb') as f: f.write(file_response.read()) print(f"已下载:{save_path}")
注意事项
- 添加
User-Agent是因为SEC服务器会拦截无标识的请求,避免被拒绝访问 os.makedirs(target_folder, exist_ok=True)确保目标文件夹存在,不会因文件夹已存在报错- 用
filing.get('type')代替直接索引,避免部分数据缺少type字段导致报错
内容的提问来源于stack exchange,提问作者joshijos
相关产品推荐
相关产品推荐

