使用Beautiful Soup提取SEC申报头信息返回None的问题求助
问题原因
SEC申报文件并非标准HTML/XML格式,Python内置的html.parser对这类非规范标记的兼容性差,无法正确识别<SEC-HEADER>这类自定义标签(标签后直接附带文本的结构不符合HTML规范),导致Beautiful Soup无法将其解析为可识别的标签节点。
解决方案
方案1:改用更宽松的解析器(lxml)
lxml对非标准标记的兼容性更好,能正确识别SEC文件中的自定义标签。先安装依赖库:
pip install lxml
修改后的代码:
import requests from bs4 import BeautifulSoup # URL of the SEC filing url = 'https://www.sec.gov/Archives/edgar/data/1166036/000110465904027382/0001104659-04-027382.txt' # Fetch the SEC filing data response = requests.get(url) content = response.content # 使用lxml解析器替代内置html.parser soup = BeautifulSoup(content, 'lxml') # 查找标签(Beautiful Soup会自动将标签名转为小写) sec_header = soup.find('sec-header') # 提取并打印头信息 if sec_header: header_data = sec_header.get_text() print(header_data) else: print("SEC-HEADER not found in the document.")
方案2:直接字符串分割(更稳定可靠)
SEC文件格式固定,直接通过字符串分割提取<SEC-HEADER>与</SEC-HEADER>之间的内容,无需依赖HTML解析器:
import requests # URL of the SEC filing url = 'https://www.sec.gov/Archives/edgar/data/1166036/000110465904027382/0001104659-04-027382.txt' # Fetch the SEC filing data response = requests.get(url) content = response.text # 定义标记并提取内容 start_marker = '<SEC-HEADER>' end_marker = '</SEC-HEADER>' if start_marker in content and end_marker in content: start_idx = content.find(start_marker) + len(start_marker) end_idx = content.find(end_marker) header_data = content[start_idx:end_idx].strip() print(header_data) else: print("SEC-HEADER not found in the document.")
内容的提问来源于stack exchange,提问作者user2946746
相关产品推荐
相关产品推荐

