You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取SEC文件Item 7.01内容:HTML标签不一致问题求助

解决SEC文件Item章节提取问题

核心问题分析

  • 原代码用精确字符串匹配"Item 7.01",但SEC文件中Item标题可能存在多个空格、制表符、换行等格式差异,导致匹配失败
  • 部分文件不存在Item 9.01,直接用它作为结束边界会导致提取内容异常

改进方案

使用正则表达式匹配Item标题,同时动态确定结束边界:

  1. 用正则匹配所有Item X.XX格式的标题,兼容任意空白字符
  2. 定位Item 7.01的起始位置后,找到后续第一个Item标题作为结束点;如果没有后续Item,则取文档末尾

修改后的代码

import urllib.request
from bs4 import BeautifulSoup
import re

url = "https://www.sec.gov/Archives/edgar/data/789019/000119312524011295/d708866d8k.htm"
headers = {'User-Agent': 'email@gmail.com'}

request2 = urllib.request.Request(url, headers=headers)
response2 = urllib.request.urlopen(request2)
html2 = response2.read()

soup2 = BeautifulSoup(html2, features="html.parser")
text2 = soup2.get_text()

# 匹配所有Item标题,兼容任意空白字符
item_pattern = re.compile(r'Item\s+\d+\.\d+')
all_items = list(item_pattern.finditer(text2))

# 定位Item 7.01的起始位置
start_match = None
for match in all_items:
    if match.group().strip() == "Item 7.01":
        start_match = match
        break

if not start_match:
    print("未找到Item 7.01")
else:
    start_idx = start_match.start()
    # 找结束位置:后续第一个Item的起始,或者文档末尾
    end_idx = len(text2)
    for match in all_items:
        if match.start() > start_idx:
            end_idx = match.start()
            break
    
    result_string = text2[start_idx:end_idx].strip()
    print(result_string)

额外优化建议

  • 可对提取的文本做进一步清理,去掉多余的换行和空白字符:result_string = re.sub(r'\s+', ' ', result_string).strip()
  • 若需更精准解析,可考虑使用SEC官方EDGAR接口或专门的Python工具库,但需严格遵守SEC的访问规则

内容的提问来源于stack exchange,提问作者user410498

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 12:01:32