从EDGAR下载10K文件时遭遇HTTP 403 Forbidden错误求助
EDGAR下载10K文件时HTTP 403错误的解决方法
出现HTTP 403 Forbidden错误的核心原因是:SEC的EDGAR服务器会拦截不带自定义User-Agent标识的请求。你用的urllib.request.urlretrieve默认没有设置请求头,会被服务器识别为非合法请求(比如爬虫),从而拒绝访问。
解决方法:给请求添加User-Agent头
由于urlretrieve无法直接设置请求头,我们需要改用urllib.request.Request对象构建请求,手动添加User-Agent信息,再通过urlopen获取文件内容后写入本地。
修改后的完整代码
import urllib.request import os import re from pathlib import Path def get_files(start_year:int, end_year:int, reform:str, inddirect:str, odirect:str): """ Downloads SEC filings for specific companies start_year -> First Year to download end_year -> Last Year to download reform -> Regex to specify forms to be downloaded inddirect -> Directory containing index files odirect -> Directory the filings will be downloaded to """ print('Downloading Filings') # 自定义User-Agent,替换成你的姓名+邮箱(SEC要求提供有效联系信息) headers = { 'User-Agent': 'Your Name your.email@example.com' } # 匹配目标文件类型的正则 re_formtype = re.compile(reform, re.IGNORECASE) # 提取文件路径的正则 re_fullfilename = re.compile(r"\|(edgar/data.*\/([\d-]+\.txt))", re.IGNORECASE) # 遍历年份 for year in range(start_year, end_year+1): download_path = os.path.join(odirect, str(year)) if not os.path.exists(download_path): os.makedirs(download_path) # 遍历季度 for qtr in range(1,5): dl_file = os.path.join(inddirect, 'master' + str(year) + str(qtr) + '.idx') # 打开已下载的索引文件 with open(dl_file, 'r') as f: count = 1 # 遍历索引文件每一行 for line in f: if count < 5: rematch = re.search(re_formtype, line) if rematch: matches = re.search(re_fullfilename, line) if matches: url = 'https://www.sec.gov/Archives/' + matches.group(1) outfile = os.path.join(download_path, matches.group(2)) if not (os.path.isfile(outfile) and os.access(outfile, os.R_OK)): print(f"Downloading: {outfile}") # 构建带请求头的请求 req = urllib.request.Request(url, headers=headers) try: with urllib.request.urlopen(req) as response, open(outfile, 'wb') as out_file: out_file.write(response.read()) count += 1 except urllib.error.HTTPError as e: print(f"下载失败 {url}: {e}") print('Downloading of Filings Complete') return # 指定要下载的文件类型(10-K相关表单) reform = '(\|10-?k(sb|sb40|405)?\s*\|)' # 索引文件存储目录 inddirect = os.path.join(Path.home(), 'edgar', 'indexfiles') # 10-K文件下载目录 odirect = os.path.join(Path.home(), 'edgar', '10K') # 执行下载函数 get_files(2018, 2019, reform, inddirect, odirect)
关键修改点说明
- 添加合法User-Agent:SEC要求请求头包含用户的联系信息(姓名+邮箱),这是避免被拦截的核心要求。
- 替换下载逻辑:用带请求头的
Request对象替代默认的urlretrieve,让服务器识别为合法请求。 - 增加异常捕获:捕获HTTP错误,方便排查单个文件的下载问题,避免程序直接崩溃。
内容的提问来源于stack exchange,提问作者Alberto Alvarez
相关产品推荐
相关产品推荐

