You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从EDGAR下载10K文件时遭遇HTTP 403 Forbidden错误求助

EDGAR下载10K文件时HTTP 403错误的解决方法

出现HTTP 403 Forbidden错误的核心原因是:SEC的EDGAR服务器会拦截不带自定义User-Agent标识的请求。你用的urllib.request.urlretrieve默认没有设置请求头,会被服务器识别为非合法请求(比如爬虫),从而拒绝访问。

解决方法:给请求添加User-Agent头

由于urlretrieve无法直接设置请求头,我们需要改用urllib.request.Request对象构建请求,手动添加User-Agent信息,再通过urlopen获取文件内容后写入本地。

修改后的完整代码

import urllib.request
import os
import re
from pathlib import Path

def get_files(start_year:int, end_year:int,
              reform:str, 
              inddirect:str, odirect:str):
    """
    Downloads SEC filings for specific companies
    start_year -> First Year to download
    end_year -> Last Year to download
    reform -> Regex to specify forms to be downloaded
    inddirect -> Directory containing index files
    odirect -> Directory the filings will be downloaded to
    """

    print('Downloading Filings')

    # 自定义User-Agent,替换成你的姓名+邮箱(SEC要求提供有效联系信息)
    headers = {
        'User-Agent': 'Your Name your.email@example.com'
    }

    # 匹配目标文件类型的正则
    re_formtype = re.compile(reform, re.IGNORECASE)
    # 提取文件路径的正则
    re_fullfilename = re.compile(r"\|(edgar/data.*\/([\d-]+\.txt))", re.IGNORECASE)

    # 遍历年份
    for year in range(start_year, end_year+1):
        download_path = os.path.join(odirect, str(year))
        if not os.path.exists(download_path):
            os.makedirs(download_path)
            
        # 遍历季度
        for qtr in range(1,5):
            dl_file = os.path.join(inddirect, 'master' + str(year) + str(qtr) + '.idx')
    
            # 打开已下载的索引文件
            with open(dl_file, 'r') as f:
                count = 1
            
                # 遍历索引文件每一行
                for line in f:
                    if count < 5:
                        rematch = re.search(re_formtype, line)
                        if rematch:
                            matches = re.search(re_fullfilename, line)
                            if matches:
                                url = 'https://www.sec.gov/Archives/' + matches.group(1)
                                outfile = os.path.join(download_path, matches.group(2))
                                
                                if not (os.path.isfile(outfile) and os.access(outfile, os.R_OK)):
                                    print(f"Downloading: {outfile}")
                                    # 构建带请求头的请求
                                    req = urllib.request.Request(url, headers=headers)
                                    try:
                                        with urllib.request.urlopen(req) as response, open(outfile, 'wb') as out_file:
                                            out_file.write(response.read())
                                        count += 1
                                    except urllib.error.HTTPError as e:
                                        print(f"下载失败 {url}: {e}")
    print('Downloading of Filings Complete')
    return
                                                   
# 指定要下载的文件类型(10-K相关表单)
reform = '(\|10-?k(sb|sb40|405)?\s*\|)'

# 索引文件存储目录
inddirect = os.path.join(Path.home(), 'edgar', 'indexfiles')

# 10-K文件下载目录
odirect = os.path.join(Path.home(), 'edgar', '10K')

# 执行下载函数
get_files(2018, 2019, reform, inddirect, odirect)

关键修改点说明

  1. 添加合法User-Agent:SEC要求请求头包含用户的联系信息(姓名+邮箱),这是避免被拦截的核心要求。
  2. 替换下载逻辑:用带请求头的Request对象替代默认的urlretrieve,让服务器识别为合法请求。
  3. 增加异常捕获:捕获HTTP错误,方便排查单个文件的下载问题,避免程序直接崩溃。

内容的提问来源于stack exchange,提问作者Alberto Alvarez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 23:40:42