如何拆分HTML字符串并提取带对应内容的各类标签
提取HTML标签及其内容的解决方案
嘿,针对你要拆分HTML字符串、提取标签和对应内容的需求,我给你两种实用的方法,根据你的场景来选就行:
方法1:正则表达式(适合简单无嵌套的HTML)
如果你的HTML结构很简单,像示例里那样都是单标签、没有嵌套的情况,用正则就能快速搞定。
JavaScript示例:
const htmlStr = `<h2>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h2> <p>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</p> <h6>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h6> <h3>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h3>`; // 匹配标签和内容的正则 const regex = /<([a-z0-9]+)>(.*?)<\/\1>/gi; const result = []; let match; while ((match = regex.exec(htmlStr)) !== null) { result.push([match[1], match[2].trim()]); } console.log(result);
运行后就能得到你期望的格式:
[ ["h2","Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum"], ["p","Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum"], ["h6","Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum"], ["h3","Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum"] ]
Python示例:
import re html_str = """<h2>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h2> <p>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</p> <h6>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h6> <h3>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h3>""" pattern = re.compile(r'<([a-z0-9]+)>(.*?)</\1>', re.IGNORECASE) result = [[match.group(1), match.group(2).strip()] for match in pattern.finditer(html_str)] print(result)
方法2:DOM解析(适合复杂HTML,更可靠)
如果你的HTML可能有嵌套标签、属性或者更复杂的结构,正则就容易出错了,这时候用DOM解析工具更稳妥。
JavaScript示例(浏览器环境或Node.js用jsdom):
// 浏览器环境直接用DOMParser const htmlStr = `<h2>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h2> <p>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</p> <h6>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h6> <h3>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h3>`; const parser = new DOMParser(); const doc = parser.parseFromString(htmlStr, 'text/html'); const elements = doc.body.children; const result = []; for (const el of elements) { result.push([el.tagName.toLowerCase(), el.textContent.trim()]); } console.log(result);
Python示例(用BeautifulSoup):
from bs4 import BeautifulSoup html_str = """<h2>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h2> <p>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</p> <h6>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h6> <h3>Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum</h3>""" soup = BeautifulSoup(html_str, 'html.parser') result = [[tag.name, tag.get_text(strip=True)] for tag in soup.body.children if tag.name] print(result)
注意:DOM解析会自动处理HTML的格式问题,比如多余的空格、换行,比正则更健壮,推荐在复杂场景下使用。
内容的提问来源于stack exchange,提问作者Shubham Nagota
相关产品推荐
相关产品推荐

