如何使用Python解析HTML标签,将指定HTML结构提取为对应格式标签列表?
实现方案
方案1:输入HTML结构规整(每行对应一个标签/内容行)
无需安装第三方库,直接对字符串按行处理即可:
# 输入的HTML字符串 html_content = '''<div class="title"> <h1> Hello World </h1> </div>''' # 按行分割,去除每行前后空白,过滤空行 result = [line.strip() for line in html_content.splitlines() if line.strip()] print(result)
运行输出:
['<div class="title">', '<h1> Hello World </h1>', '</div>']
方案2:兼容任意排版的HTML
如果输入的HTML是压缩成单行、缩进混乱的格式,建议用BeautifulSoup先做格式化再处理,稳定性更高:
- 先安装依赖:
pip install beautifulsoup4
- 实现代码:
from bs4 import BeautifulSoup # 就算输入是压缩后的HTML也能正常处理 html_content = '<div class="title"><h1> Hello World </h1></div>' soup = BeautifulSoup(html_content, 'html.parser') # 先将HTML格式化为每行一个节点的结构 formatted_html = soup.prettify() result = [line.strip() for line in formatted_html.splitlines() if line.strip()] print(result)
运行输出和方案1完全一致。
内容的提问来源于stack exchange,提问作者MyDisplay
相关产品推荐
相关产品推荐

