You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R爬取SEC Edgar网站遇HTTP 403错误求助

解决SEC Edgar爬取时的HTTP 403错误

HTTP 403错误是因为SEC Edgar服务器拒绝了你的请求——SEC要求所有爬取请求必须携带包含有效联系方式的User-Agent头,默认的R请求头会被拦截。以下是具体解决方案:

1. 添加合法的User-Agent请求头

用httr包发送请求(推荐)

httr包可以灵活设置请求头,示例代码:

library(httr)
library(qdapRegex) # 假设你的rm_between来自这个包

# 替换成你的姓名和邮箱
headers <- add_headers(
  `User-Agent` = "张三 zhangsan@example.com",
  `Accept-Encoding` = "gzip, deflate"
)

# 目标URL
target_url <- "https://www.sec.gov/Archives/edgar/data/0001089892/000156821423000002/0001568214-23-000002.txt"

# 发送请求
response <- GET(target_url, headers)

# 处理响应
if (status_code(response) == 200) {
  raw_text <- content(response, "text", encoding = "UTF-8")
  # 调用你的自定义函数
  result <- getRawValue("your_parent_tag", raw_text)
} else {
  stop(paste("请求失败,状态码:", status_code(response)))
}

用base R的readLines

如果坚持用base R,直接设置useragent参数:

target_url <- "https://www.sec.gov/Archives/edgar/data/0001089892/000156821423000002/0001568214-23-000002.txt"
raw_text <- readLines(target_url, warn = FALSE,
                      useragent = "张三 zhangsan@example.com")
raw_text <- paste(raw_text, collapse = "\n")

2. 优化你的getRawValue函数

手动处理字符串解析XML容易出错,建议用专业的XML解析包(比如rvest或XML)替代:

library(rvest)

getRawValue <- function(parentTag, rawText) {
  # 解析XML文档
  xml_doc <- read_xml(rawText)
  # 用XPath定位父标签下的所有<value>元素并提取文本
  xml_nodes(xml_doc, paste0("//", parentTag, "/value")) %>% 
    xml_text()
}

3. 爬取注意事项

  • 必须使用真实的联系方式作为User-Agent,SEC会验证该信息,虚假信息可能导致IP被封禁
  • 批量爬取时添加延迟(比如Sys.sleep(1)),遵守SEC的爬虫速率限制
  • 避免频繁重复请求同一资源

内容的提问来源于stack exchange,提问作者js80

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 19:35:45