You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分HTML字符串并提取带对应内容的各类标签

提取HTML标签及其内容的解决方案

嘿,针对你要拆分HTML字符串、提取标签和对应内容的需求,我给你两种实用的方法,根据你的场景来选就行:

方法1:正则表达式(适合简单无嵌套的HTML)

如果你的HTML结构很简单,像示例里那样都是单标签、没有嵌套的情况,用正则就能快速搞定。

JavaScript示例:

const htmlStr = `<h2>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h2> <p>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</p> <h6>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h6> <h3>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h3>`;

// 匹配标签和内容的正则
const regex = /<([a-z0-9]+)>(.*?)<\/\1>/gi;
const result = [];
let match;

while ((match = regex.exec(htmlStr)) !== null) {
  result.push([match[1], match[2].trim()]);
}

console.log(result);

运行后就能得到你期望的格式:

[
  ["h2","Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum"],
  ["p","Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum"],
  ["h6","Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum"],
  ["h3","Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum"]
]

Python示例:

import re

html_str = """<h2>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h2> <p>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</p> <h6>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h6> <h3>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h3>"""

pattern = re.compile(r'<([a-z0-9]+)>(.*?)</\1>', re.IGNORECASE)
result = [[match.group(1), match.group(2).strip()] for match in pattern.finditer(html_str)]

print(result)

方法2:DOM解析(适合复杂HTML,更可靠)

如果你的HTML可能有嵌套标签、属性或者更复杂的结构,正则就容易出错了,这时候用DOM解析工具更稳妥。

JavaScript示例(浏览器环境或Node.js用jsdom):

// 浏览器环境直接用DOMParser
const htmlStr = `<h2>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h2> <p>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</p> <h6>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h6> <h3>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h3>`;

const parser = new DOMParser();
const doc = parser.parseFromString(htmlStr, 'text/html');
const elements = doc.body.children;
const result = [];

for (const el of elements) {
  result.push([el.tagName.toLowerCase(), el.textContent.trim()]);
}

console.log(result);

Python示例(用BeautifulSoup):

from bs4 import BeautifulSoup

html_str = """<h2>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h2> <p>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</p> <h6>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h6> <h3>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h3>"""

soup = BeautifulSoup(html_str, 'html.parser')
result = [[tag.name, tag.get_text(strip=True)] for tag in soup.body.children if tag.name]

print(result)

注意:DOM解析会自动处理HTML的格式问题,比如多余的空格、换行,比正则更健壮,推荐在复杂场景下使用。

内容的提问来源于stack exchange,提问作者Shubham Nagota

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 07:42:42