如何从HTML字符串的h1标签提取纯文本?支持jQuery实现
解决方法:提取HTML字符串中H1标签的纯文本
纯JavaScript正则方案(改进版)
你当前的正则/<h1>(.*?)<\/h1>/g搭配match()会返回完整的H1标签(比如<h1>topic 1</h1>),这是因为带g标志的match()只返回匹配到的完整字符串,不会捕获分组内容。要拿到标签内的文本,可以用以下两种方式:
方法1:使用matchAll()捕获分组
matchAll()会返回包含所有匹配结果(包括分组)的迭代器,我们可以将其转为数组后提取分组内容:
let str = `<h1>topic 1</h1><p>desc of topic 1</p><h1>topic 2</h1><p>desc of topic 2</p>`; // 用matchAll获取所有匹配项,提取每个匹配的第1个分组(即标签内文本) const h1Texts = Array.from(str.matchAll(/<h1>(.*?)<\/h1>/g), match => match[1]); console.log(h1Texts); // 输出: ["topic 1", "topic 2"]
方法2:使用exec()循环匹配
如果你的环境不支持matchAll()(比如旧版浏览器),可以用exec()配合循环来逐个捕获分组:
let str = `<h1>topic 1</h1><p>desc of topic 1</p><h1>topic 2</h1><p>desc of topic 2</p>`; const regex = /<h1>(.*?)<\/h1>/g; let match; const h1Texts = []; // 循环执行exec,直到没有匹配项 while ((match = regex.exec(str)) !== null) { h1Texts.push(match[1]); // 提取分组内容 } console.log(h1Texts); // 输出: ["topic 1", "topic 2"]
jQuery方案(更可靠,推荐)
正则处理HTML容易遇到边界情况(比如H1标签带属性:<h1 class="title">Hello</h1>,或者标签内有换行),用jQuery的DOM解析能力会更稳定:
let str = `<h1>topic 1</h1><p>desc of topic 1</p><h1>topic 2</h1><p>desc of topic 2</p>`; // 将字符串转为jQuery对象,遍历H1标签提取文本 const h1Texts = $(str).find('h1').map(function() { return $(this).text(); }).get(); // 将jQuery对象转为普通数组 console.log(h1Texts); // 输出: ["topic 1", "topic 2"]
这种方法不需要关心H1标签的格式细节,jQuery会自动处理DOM结构,兼容性和鲁棒性更好。
内容的提问来源于stack exchange,提问作者totalnoob
相关产品推荐
相关产品推荐

