You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup提取SEC申报头信息返回None的问题求助

问题原因

SEC申报文件并非标准HTML/XML格式,Python内置的html.parser对这类非规范标记的兼容性差,无法正确识别<SEC-HEADER>这类自定义标签(标签后直接附带文本的结构不符合HTML规范),导致Beautiful Soup无法将其解析为可识别的标签节点。

解决方案

方案1:改用更宽松的解析器(lxml)

lxml对非标准标记的兼容性更好,能正确识别SEC文件中的自定义标签。先安装依赖库:

pip install lxml

修改后的代码:

import requests
from bs4 import BeautifulSoup

# URL of the SEC filing
url = 'https://www.sec.gov/Archives/edgar/data/1166036/000110465904027382/0001104659-04-027382.txt'

# Fetch the SEC filing data
response = requests.get(url)
content = response.content

# 使用lxml解析器替代内置html.parser
soup = BeautifulSoup(content, 'lxml')

# 查找标签(Beautiful Soup会自动将标签名转为小写)
sec_header = soup.find('sec-header')

# 提取并打印头信息
if sec_header:
    header_data = sec_header.get_text()
    print(header_data)
else:
    print("SEC-HEADER not found in the document.")

方案2:直接字符串分割(更稳定可靠)

SEC文件格式固定,直接通过字符串分割提取<SEC-HEADER>与</SEC-HEADER>之间的内容,无需依赖HTML解析器:

import requests

# URL of the SEC filing
url = 'https://www.sec.gov/Archives/edgar/data/1166036/000110465904027382/0001104659-04-027382.txt'

# Fetch the SEC filing data
response = requests.get(url)
content = response.text

# 定义标记并提取内容
start_marker = '<SEC-HEADER>'
end_marker = '</SEC-HEADER>'

if start_marker in content and end_marker in content:
    start_idx = content.find(start_marker) + len(start_marker)
    end_idx = content.find(end_marker)
    header_data = content[start_idx:end_idx].strip()
    print(header_data)
else:
    print("SEC-HEADER not found in the document.")

内容的提问来源于stack exchange,提问作者user2946746

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 21:23:15