You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取HTML页面中<h3>与</h3>标签间的文本?

Python提取HTML中

标签内文本的方法

方法一:使用BeautifulSoup(推荐)

对于HTML解析,专用解析库比正则表达式更可靠,能处理复杂的HTML结构。

  1. 先安装依赖:
pip install beautifulsoup4
  1. 编写代码:
from bs4 import BeautifulSoup

html_content = """
<h3 id="basics">1. Creating a Web Page</h3>
<p>
Once you've made your "home page" (index.html) you can add more pages to
your site, and your home page can link to them.
<h3 id="syntax">>2. HTML Syntax</h3>
"""

# 解析HTML内容
soup = BeautifulSoup(html_content, 'html.parser')
# 提取所有<h3>标签的文本,strip=True去除首尾空白
h3_texts = [h3.get_text(strip=True) for h3 in soup.find_all('h3')]

print(h3_texts)
# 输出结果: ['1. Creating a Web Page', '>2. HTML Syntax']

方法二:使用正则表达式(适合简单场景)

如果HTML结构简单且固定,可以用正则快速匹配,但不推荐用于复杂HTML(比如标签嵌套、属性多变的情况)。

代码示例:

import re

html_content = """
<h3 id="basics">1. Creating a Web Page</h3>
<p>
Once you've made your "home page" (index.html) you can add more pages to
your site, and your home page can link to them.
<h3 id="syntax">>2. HTML Syntax</h3>
"""

# 正则模式:匹配带任意属性的<h3>标签,提取中间文本
pattern = r'<h3.*?>(.*?)</h3>'
# re.DOTALL让.匹配换行符,处理标签内文本换行的情况
matches = re.findall(pattern, html_content, re.DOTALL)
# 清理结果中的首尾空白
h3_texts = [text.strip() for text in matches]

print(h3_texts)
# 输出结果: ['1. Creating a Web Page', '>2. HTML Syntax']

两种方法对比

  • BeautifulSoup:容错性强,能处理不规范的HTML,支持复杂结构解析,是HTML处理的首选方案。
  • 正则表达式:代码更简洁,但对HTML结构变化的适应性差,若遇到标签嵌套、属性格式变化等情况容易匹配失败。

内容的提问来源于stack exchange,提问作者Xazratxon Sultonxonov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 02:35:16