You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python实现标签查找器 提取标记代码中的所有HTML标签

Python提取HTML标签的实现方法

现有待处理的标记代码片段如下:

<h1>hello
<p>world
<b>from
<li>python

需要提取所有使用的标签,输出格式为['h1', 'p', 'b', 'li'],可通过以下两种常用方法实现:

方法1:正则表达式匹配(适合简单规整的文本场景)

无需额外安装依赖,直接通过正则规则匹配<>包裹的标签名:

import re

html_content = """<h1>hello
<p>world
<b>from
<li>python"""

# 匹配<开头后紧跟的单词字符作为标签名
tags = re.findall(r'<(\w+)>', html_content)
print(tags)  # 输出: ['h1', 'p', 'b', 'li']

注意:该方法仅适合格式简单、无标签属性、无特殊嵌套场景的HTML片段,复杂场景下容易出现匹配误差。

方法2:使用BeautifulSoup HTML解析库(推荐,适合所有HTML场景)

调用专业的HTML解析库处理,容错性更高,即使标签带属性、格式不完整也能正确提取:

  1. 先安装依赖库:
pip install beautifulsoup4
  1. 实现代码:
from bs4 import BeautifulSoup

html_content = """<h1>hello
<p>world
<b>from
<li>python"""

soup = BeautifulSoup(html_content, 'html.parser')
# 遍历所有解析到的标签,提取标签名
tags = [tag.name for tag in soup.find_all()]
print(tags)  # 输出: ['h1', 'p', 'b', 'li']

内容的提问来源于stack exchange,提问作者user17231840

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 20:36:00