如何使用BeautifulSoup抓取每次变化的Central Index Key?
用BeautifulSoup提取XML文档中的Central Index Key(CIK)
针对你要提取SEC文档中动态变化的Central Index Key(CIK)需求,这里提供两种可靠的实现方式,同时可以将每个文档的唯一标识与对应CIK关联记录:
方法1:遍历文本行提取
由于SEC文档的头部内容多为结构化文本而非严格嵌套的XML节点,直接遍历文本行定位目标字段更直观:
from bs4 import BeautifulSoup # 示例:读取文档内容(实际场景可从文件/网络请求获取) with open("sec_document.xml", "r") as f: xml_content = f.read() soup = BeautifulSoup(xml_content, "lxml") # 定位SEC-HEADER标签 sec_header = soup.find("sec-header") if not sec_header: print("未找到SEC-HEADER标签") exit() # 按行分割头部文本 header_lines = sec_header.get_text().split("\n") target_cik = None for line in header_lines: if "CENTRAL INDEX KEY:" in line: # 分割字段名与值,去除多余空格 target_cik = line.split("CENTRAL INDEX KEY:")[1].strip() break # 输出结果 if target_cik: print(f"提取到的CIK: {target_cik}") else: print("文档中未找到CIK")
方法2:正则表达式匹配
用正则表达式可以更简洁地匹配CIK字段,适合处理格式固定的文本:
import re from bs4 import BeautifulSoup xml_content = # 你的文档内容 soup = BeautifulSoup(xml_content, "lxml") sec_header_text = soup.find("sec-header").get_text() # 匹配CENTRAL INDEX KEY后的数字串 match_result = re.search(r"CENTRAL INDEX KEY:\s+(\d+)", sec_header_text) if match_result: target_cik = match_result.group(1) print(f"提取到的CIK: {target_cik}") else: print("未找到有效CIK")
记录多文档的CIK对应关系
如果需要批量处理多个文档,只需将每个文档的唯一标识(比如文件名、accession number)与提取到的CIK关联存储,示例如下:
import re from bs4 import BeautifulSoup # 假设docs是包含多个文档的列表,每个元素为(文档唯一ID, 文档内容) document_cik_mapping = {} for doc_id, content in docs: soup = BeautifulSoup(content, "lxml") sec_header = soup.find("sec-header") if not sec_header: document_cik_mapping[doc_id] = None continue match = re.search(r"CENTRAL INDEX KEY:\s+(\d+)", sec_header.get_text()) document_cik_mapping[doc_id] = match.group(1) if match else None # 查看所有文档的CIK记录 for doc_id, cik in document_cik_mapping.items(): print(f"文档ID {doc_id} 对应的CIK: {cik}")
内容的提问来源于stack exchange,提问作者pilotso
相关产品推荐
相关产品推荐

