You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python+Selenium提取SEC归档网站各类Filer状态的通用代码咨询

通用提取SEC归档中公司分类信息的解决方案

面对SEC归档里五花八门的HTML结构,别纠结标签嵌套了——核心思路是抓关键词和它紧邻的勾选符号,不管HTML怎么变,这俩的对应关系是不会乱的。下面是具体的实现方案:

核心思路

不管是哪种HTML结构,我们要提取的5类公司标识(Large accelerated filer、Accelerated filer等)和它的勾选状态(选中/未选中/未提及),本质是「关键词 + 相邻符号」的配对。所以我们可以:

  1. 先把HTML里的干扰项(比如 、多余标签、换行)清理干净,转成易处理的文本格式
  2. 针对每个目标关键词,找到它后面最近的勾选符号(常见的选中符号:☒、x、þ;未选中符号:☐、¨)
  3. 处理缺失的选项:如果某个关键词在文本里找不到,直接标记为「未提及」

Python代码实现

这里用BeautifulSoup做HTML解析,再用正则匹配关键词和符号的对应关系:

from bs4 import BeautifulSoup
import re

def extract_filer_status(html_content):
    # 1. 预处理HTML:清理特殊字符和多余标签
    soup = BeautifulSoup(html_content, "html.parser")
    # 提取所有文本,替换 为空格,去掉多余空白
    clean_text = soup.get_text().replace("\xa0", " ").strip()
    # 把连续空白换成单个空格
    clean_text = re.sub(r"\s+", " ", clean_text)

    # 2. 定义目标关键词(包含可能的变体,比如大小写差异)
    target_filers = {
        "Large accelerated filer": re.compile(r"Large\s+accelerated\s+filer", re.IGNORECASE),
        "Accelerated filer": re.compile(r"Accelerated\s+filer", re.IGNORECASE),
        "Non-accelerated filer": re.compile(r"Non-accelerated\s+filer", re.IGNORECASE),
        "Smaller reporting company": re.compile(r"Smaller\s+reporting\s+(company|Company)", re.IGNORECASE),
        "Emerging growth company": re.compile(r"Emerging\s+growth\s+company", re.IGNORECASE)
    }

    # 3. 定义勾选符号的规则:选中/未选中
    check_symbols = {
        "selected": ["☒", "x", "þ"],
        "unselected": ["☐", "¨"]
    }

    result = {}

    for filer_name, pattern in target_filers.items():
        # 找到关键词在文本中的位置
        match = pattern.search(clean_text)
        if not match:
            result[filer_name] = "未提及"
            continue
        
        # 从关键词结束的位置开始,找最近的符号
        start_pos = match.end()
        # 截取关键词后面的一段文本,找第一个符号
        sub_text = clean_text[start_pos:]
        # 匹配所有可能的符号
        symbol_match = re.search(r"([☒xþ☐¨])", sub_text)
        
        if not symbol_match:
            result[filer_name] = "无状态标记"
            continue
        
        symbol = symbol_match.group(1)
        if symbol in check_symbols["selected"]:
            result[filer_name] = "是"
        elif symbol in check_symbols["unselected"]:
            result[filer_name] = "否"
        else:
            result[filer_name] = "未知符号"
    
    return result

# 测试示例(用你提供的三种结构)
if __name__ == "__main__":
    # 结构1的HTML
    html1 = """<td valign="bottom">Large&nbsp;accelerated&nbsp;filer</td> <td valign="bottom">&nbsp;</td> <td valign="bottom">☒</td> <td valign="bottom">&nbsp;&nbsp;</td> <td valign="bottom">Accelerated&nbsp;filer</td> <td valign="bottom">&nbsp;</td> <td valign="bottom">☐</td></tr> <tr style="page-break-inside:avoid ; font-family:Times New Roman; font-size:10pt"> <td valign="bottom"><font style="white-space:nowrap">Non-accelerated&nbsp;filer</font></td> <td valign="bottom">&nbsp;</td> <td valign="bottom">☐&nbsp;&nbsp;(Do not check if a smaller reporting company)</td> <td valign="bottom">&nbsp;&nbsp;</td> <td valign="bottom">Smaller&nbsp;reporting&nbsp;company</td> <td valign="bottom">&nbsp;</td> <td valign="bottom">☐</td></tr> <tr style="page-break-inside:avoid ; font-family:Times New Roman; font-size:10pt"> <td valign="bottom">Emerging&nbsp;growth&nbsp;company</td> <td valign="bottom">&nbsp;</td> <td valign="bottom">☐</td> <td valign="bottom">&nbsp;&nbsp;</td> <td valign="bottom"></td> <td valign="bottom">&nbsp;</td> <td valign="bottom"></td></tr>"""
    print("结构1提取结果:", extract_filer_status(html1))

    # 结构2的HTML
    html2 = """filer&nbsp;&nbsp;<font style="FONT-FAMILY:WINGDINGS">x</font>&nbsp;&nbsp;&nbsp;&nbsp;Accelerated filer&nbsp;&nbsp;<font style="FONT-FAMILY:WINGDINGS">¨</font>&nbsp;&nbsp;&nbsp;&nbsp;Non-accelerated filer&nbsp;&nbsp;<font style="FONT-FAMILY:WINGDINGS">¨</font>&nbsp;&nbsp;&nbsp;&nbsp;Smaller reporting company&nbsp;&nbsp;<font style="FONT-FAMILY:WINGDINGS">¨</font> </font>"""
    print("结构2提取结果:", extract_filer_status(html2))

    # 结构3的HTML
    html3 = """<tbody><tr> <td width="63%"></td> <td valign="bottom" width="2%"></td> <td width="35%"></td></tr> <tr> <td valign="top"> <p style="text-indent:2.00em"><font face="Times New Roman" size="2">Large accelerated filer&nbsp;&nbsp;<font face="WINGDINGS">¨</font></font></p></td> <td valign="bottom"><font size="1">&nbsp;&nbsp;</font></td> <td valign="bottom"><font face="Times New Roman" size="2">Accelerated filer&nbsp;&nbsp;<font face="WINGDINGS">þ</font></font></td></tr> <tr> <td valign="top"> <p style="text-indent:2.00em"><font face="Times New Roman" size="2">Non-accelerated filer&nbsp;&nbsp;<font face="WINGDINGS">¨</font>&nbsp;&nbsp; (Do not check if a smaller reporting company)</font></p></td> <td valign="bottom"><font size="1">&nbsp;&nbsp;</font></td> <td valign="bottom"><font face="Times New Roman" size="2">Smaller reporting Company&nbsp;&nbsp;<font face="WINGDINGS">¨</font></font></td></tr> </tbody>"""
    print("结构3提取结果:", extract_filer_status(html3))

代码说明

  • 预处理步骤:用BeautifulSoup提取纯文本,清理掉HTML里的特殊空格和冗余内容,避免标签干扰
  • 正则匹配关键词:用忽略大小写的正则,兼容"Smaller reporting Company"这种大小写变体
  • 符号匹配:只找关键词后面第一个出现的勾选符号,确保对应关系准确
  • 容错处理:如果关键词没找到或者没有符号,都会给出明确的标记,不会报错

这个方案能适配你遇到的三种结构,甚至以后遇到新的结构,只要关键词和符号的对应关系不变,就能正常提取。

内容的提问来源于stack exchange,提问作者PriyankaJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 22:22:30