You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用JavaScript正则表达式提取HTML标签中的文本为数组

Extract Text Inside Escaped HTML Tags into an Array

Got it, let's work through this. Your current regex just strips all escaped HTML tags and glues the text into a single string, but we need to capture each tag's content separately into an array. Here's a straightforward, reliable approach:

Step-by-Step Solution

First, we have to decode the escaped HTML entities (< and >) into actual angle brackets—this lets our regex properly identify opening and closing tags. Then we can extract the text inside each tag pair.

const params = '<b>ppshein</b><span>male</span><u>javascript</u>';

// 1. Decode escaped HTML entities to real opening/closing tags
const decodedHtml = params.replace(/&lt;/g, '<').replace(/&gt;/g, '>');

// 2. Extract text inside each tag using regex with lookarounds
const extractedTextArray = decodedHtml.match(/(?<=<[^>]+>)([^<]+)(?=<\/[^>]+)/g) || [];

console.log(extractedTextArray); // Output: ['ppshein', 'male', 'javascript']

How the Regex Works

Let’s break down the pattern /(?<=<[^>]+>)([^<]+)(?=<\/[^>]+)/g:

  • (?<=<[^>]+>): Positive lookbehind – makes sure we’re positioned right after an opening tag (starts with <, ends with >)
  • ([^<]+): Capture group – grabs all characters that aren’t < (this is the actual text inside the tag)
  • (?=<\/[^>]+): Positive lookahead – ensures we’re positioned right before a closing tag (starts with </, ends with >)
  • The g flag tells the regex to find all matches instead of just the first one.

Edge Case Handling

If there’s a chance some tags might be empty (like &lt;b&gt;&lt;/b&gt;), add a quick filter to remove empty strings from the final array:

const cleanedArray = extractedTextArray.filter(text => text.trim() !== '');

内容的提问来源于stack exchange,提问作者Pyae Phyoe Shein

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 13:27:50