You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用JavaScript处理API返回的HTML,提取<a>标签内文本?

解决方案

你的问题出在正则表达式的锚点使用和匹配逻辑上:

  • 正则里的^和$是行首/行尾锚定,加上gm模式后,只会匹配整行完全是<a>标签的内容,但你的目标文本里,<a>标签是嵌入在普通文本中的,所以根本无法匹配到目标内容。
  • 另外,.*的贪婪匹配可能会导致意外的跨标签匹配,不够严谨。

方法1:修正正则表达式

把锚点去掉,改用更精准的全局匹配规则,专门捕获<a>标签内的文本内容:

const r = /<a[^>]*>(.*?)<\/a>/g;
let link = `<a href="https://google.com" target="_blank">google.com</a> test <a href="test.com">test.com</a>`;

link = link.replaceAll(r, "$1");
console.log(link);
// 输出:google.com test test.com

正则说明:

  • <a[^>]*>:匹配<a开头,后面跟着任意非>的字符(即所有a标签的属性),直到>结束
  • (.*?):非贪婪捕获a标签内的文本内容
  • <\/a>:匹配闭合标签
  • g:全局匹配所有符合条件的a标签

方法2:用DOM API处理(更可靠)

正则处理HTML容易出现边缘情况(比如标签内有特殊字符、嵌套标签等),用浏览器原生的DOM解析器更稳妥:

function htmlToPlainText(html) {
  const parser = new DOMParser();
  const doc = parser.parseFromString(html, "text/html");
  return doc.body.textContent || "";
}

const htmlContent = `Lorem ipsum dolor sit amet <a href="https://example.com">example.com</a>
Pellentesque porta ligula et justo condimentum, nec tincidunt libero tempor.
Pellentesque nunc justo, tincidunt sit amet suscipit sit amet, auctor <a href="https://google.com">google.com</a>`;

const plainText = htmlToPlainText(htmlContent);
console.log(plainText);
// 输出你想要的纯文本格式

这个方法会自动处理所有HTML标签,只保留文本内容,完全避免正则的局限性。

内容的提问来源于stack exchange,提问作者ppoz21

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 02:06:24