技术问询:如何从SEC 10-K HTML文件中提取仅可见文本
提取SEC 10-K文件中的可见文本(移除表格、XBRL等内容)
我用edgar包从SEC EDGAR系统下载了10-K文件,每份文件的底层HTML代码存储在单独文本文件中。需求是仅提取HTML页面读者可见的文本,移除所有表格、图形、XBRL、Zip文档等内容。
现有下载代码
# 加载包 require(edgar) # SEC公司标识符(3家公司) cik_codes = c(0000099780, 0000010456, 0000099780) # 时间范围(每家公司10份报告) years = c(2011:2020) # 下载文件 getFilings(cik.no = cik_codes, form.type = "10-K", filing.year = years, downl.permit = "y", useragent = "Name Surname E-Mail")
当前处理代码(存在残留问题)
# 加载包 require(XML) require(readr) # 读取文件 filing = readr::read_file(path) # 解析文档 doc = XML::htmlParse(filing, asText = TRUE, useInternalNodes = TRUE, addFinalizer = FALSE) # 提取标签外文本(仍残留表格、XBRL等内容) f.text = XML::xpathSApply(doc, "//text()[not(ancestor::script)][not(ancestor::style)][not(ancestor::noscript)][not(ancestor::form)]", XML::xmlValue)
改进方案
当前XPath未排除表格、图形、XBRL相关节点,需要扩展过滤规则,同时清理空白文本:
# 加载包 require(XML) require(readr) require(stringr) # 读取文件 filing = readr::read_file(path) # 解析文档 doc = XML::htmlParse(filing, asText = TRUE, useInternalNodes = TRUE, addFinalizer = FALSE) # 扩展XPath,排除更多不需要的节点,同时过滤空白文本 f.text = XML::xpathSApply(doc, "//text()[ not(ancestor::script) and not(ancestor::style) and not(ancestor::noscript) and not(ancestor::form) and not(ancestor::table) and not(ancestor::img) and not(ancestor::object) and not(ancestor::svg) and not(ancestor::iframe) ]", XML::xmlValue) # 清理空白:移除纯空白字符串,压缩连续空白 clean_text = f.text[str_trim(f.text) != ""] clean_text = str_squish(clean_text) # 合并为完整文本 final_text = paste(clean_text, collapse = " ")
关键改进点:
- 新增排除规则:
ancestor::table(表格)、ancestor::img(图片)、ancestor::object(XBRL/Zip文档通常嵌入在此标签)、ancestor::svg(矢量图形)、ancestor::iframe(内嵌框架) - 使用
stringr包清理空白文本:移除纯空行/空白字符串,压缩连续空格、换行符等 - 最后将所有有效文本合并为完整字符串
内容的提问来源于stack exchange,提问作者YPG
相关产品推荐
相关产品推荐

