You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何仅使用urllib库获取SEC指定网页?解决403访问受限问题

解决urllib请求SEC官网403 Forbidden错误的方法

SEC的反爬机制不只是检查User-Agent,单一添加UA不足以通过验证,你需要模拟更完整的浏览器请求头来绕过403限制。

具体实现代码

import urllib.request

target_url = "https://www.sec.gov/files/company_tickers_exchange.json"

# 模拟真实Chrome浏览器的请求头
request_headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Accept-Encoding': 'gzip, deflate, br',
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
}

# 构造请求并发送
req = urllib.request.Request(target_url, headers=request_headers)
try:
    with urllib.request.urlopen(req) as resp:
        # 读取并解码返回内容
        content = resp.read().decode('utf-8')
        print(content[:300])  # 打印部分内容验证请求成功
except urllib.error.HTTPError as err:
    print(f"请求失败: {err}")

关键说明

  • 补充的Accept、Accept-Language等头信息是浏览器请求的标准字段,能让请求更贴近真实用户行为
  • 可以根据当前主流浏览器版本更新User-Agent的值,避免因UA过时被拦截
  • 注意控制请求频率,遵守SEC网站的访问规则,防止IP被封禁

内容的提问来源于stack exchange,提问作者showkey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 13:53:20