使用CIK编号获取公司最新10-K文件URL遇HTTP 403错误求助
解决SEC Edgar请求403错误的方案
SEC的Edgar服务器会拦截未携带合法标识的请求,你的代码出现403是因为默认requests请求头被识别为非合规爬虫请求。解决核心是添加符合要求的请求头,尤其是User-Agent字段(SEC要求该字段包含你的联系信息,比如邮箱,避免被封禁)。
修改后的代码如下:
import requests from bs4 import BeautifulSoup # CIK number for Apple is 0001166559 cik_number = "0001166559" url = f"https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK={cik_number}&type=10-K&dateb=&owner=exclude&count=40" # 替换成你的姓名和邮箱,合规标识请求身份 headers = { 'User-Agent': 'Your Full Name <your.email@example.com>', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8' } response = requests.get(url, headers=headers) # 主动抛出请求错误,方便排查问题 response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') # 定位最新10-K文件的文档按钮 link = soup.find('a', {'id': 'documentsbutton'}) if link: # 拼接完整的SEC域名,原href是相对路径 filing_url = f"https://www.sec.gov{link['href']}" print(filing_url) else: print("未找到最新的10-K文件链接")
额外注意事项:
- 严格遵守SEC爬虫规则,不要高频发送请求,建议添加请求间隔
- 处理
link为None的情况,避免抛出KeyError - 若后续需要批量爬取,建议参考SEC官方的爬虫指南控制请求频率
内容的提问来源于stack exchange,提问作者Sushmitha Krishnan
相关产品推荐
相关产品推荐

