Python提取SEC文件Item 7.01内容:HTML标签不一致问题求助
解决SEC文件Item章节提取问题
核心问题分析
- 原代码用精确字符串匹配
"Item 7.01",但SEC文件中Item标题可能存在多个空格、制表符、换行等格式差异,导致匹配失败 - 部分文件不存在
Item 9.01,直接用它作为结束边界会导致提取内容异常
改进方案
使用正则表达式匹配Item标题,同时动态确定结束边界:
- 用正则匹配所有
Item X.XX格式的标题,兼容任意空白字符 - 定位
Item 7.01的起始位置后,找到后续第一个Item标题作为结束点;如果没有后续Item,则取文档末尾
修改后的代码
import urllib.request from bs4 import BeautifulSoup import re url = "https://www.sec.gov/Archives/edgar/data/789019/000119312524011295/d708866d8k.htm" headers = {'User-Agent': 'email@gmail.com'} request2 = urllib.request.Request(url, headers=headers) response2 = urllib.request.urlopen(request2) html2 = response2.read() soup2 = BeautifulSoup(html2, features="html.parser") text2 = soup2.get_text() # 匹配所有Item标题,兼容任意空白字符 item_pattern = re.compile(r'Item\s+\d+\.\d+') all_items = list(item_pattern.finditer(text2)) # 定位Item 7.01的起始位置 start_match = None for match in all_items: if match.group().strip() == "Item 7.01": start_match = match break if not start_match: print("未找到Item 7.01") else: start_idx = start_match.start() # 找结束位置:后续第一个Item的起始,或者文档末尾 end_idx = len(text2) for match in all_items: if match.start() > start_idx: end_idx = match.start() break result_string = text2[start_idx:end_idx].strip() print(result_string)
额外优化建议
- 可对提取的文本做进一步清理,去掉多余的换行和空白字符:
result_string = re.sub(r'\s+', ' ', result_string).strip() - 若需更精准解析,可考虑使用SEC官方EDGAR接口或专门的Python工具库,但需严格遵守SEC的访问规则
内容的提问来源于stack exchange,提问作者user410498
相关产品推荐
相关产品推荐

