如何用Python从HTML文件中提取所有嵌套文本(尽量不依赖库)
提取HTML中所有文本的两种方案
方案一:纯正则实现(无第三方库)
如果不想引入第三方库,可以用正则匹配并移除所有HTML标签,再清理多余空白来提取文本。核心思路是把所有<>包裹的标签内容替换为空,剩下的就是文本内容。
示例代码:
import re html_content = """<div> <div> <h1>Title</h1> <p>Some text i want to fetch</p> </div> <RandomVueComponent>Some other text i want to fetch</RandomVueComponent> </div> <a href="#">Might be some text too</a> """ # 移除所有HTML标签 clean_text = re.sub(r'<[^>]+>', '', html_content) # 清理多余的换行、空格,过滤空行 text_lines = [line.strip() for line in clean_text.split('\n') if line.strip()] # 合并成最终文本(也可以保留列表形式按行输出) final_text = '\n'.join(text_lines) print(final_text)
输出结果:
Title Some text i want to fetch Some other text i want to fetch Might be some text too
注意:正则处理HTML有局限性,如果HTML里存在标签属性包含>(比如<div data-x="a>b">)、注释、<script>/<style>这类不需要提取的文本块,正则会失效。这种情况下更推荐用下面的方案。
方案二:用BeautifulSoup(无需指定标签)
你之前对BeautifulSoup的理解有误——它完全不需要以特定标签作为匹配依据,直接调用get_text()就能提取所有嵌套层级的文本,不管标签类型。而且它能正确处理复杂HTML结构,避免正则的各种坑。
示例代码:
from bs4 import BeautifulSoup html_content = """<div> <div> <h1>Title</h1> <p>Some text i want to fetch</p> </div> <RandomVueComponent>Some other text i want to fetch</RandomVueComponent> </div> <a href="#">Might be some text too</a> """ soup = BeautifulSoup(html_content, 'html.parser') # 提取所有文本,参数strip=True会自动清理多余空白,separator指定分隔符 clean_text = soup.get_text(strip=True, separator='\n') print(clean_text)
输出结果和正则方案一致,但能处理更复杂的HTML场景,比如包含注释、特殊属性的标签等。
内容的提问来源于stack exchange,提问作者Pous17
相关产品推荐
相关产品推荐

