如何用JavaScript正则表达式提取HTML标签中的文本为数组
Got it, let's work through this. Your current regex just strips all escaped HTML tags and glues the text into a single string, but we need to capture each tag's content separately into an array. Here's a straightforward, reliable approach:
Step-by-Step Solution
First, we have to decode the escaped HTML entities (< and >) into actual angle brackets—this lets our regex properly identify opening and closing tags. Then we can extract the text inside each tag pair.
const params = '<b>ppshein</b><span>male</span><u>javascript</u>'; // 1. Decode escaped HTML entities to real opening/closing tags const decodedHtml = params.replace(/</g, '<').replace(/>/g, '>'); // 2. Extract text inside each tag using regex with lookarounds const extractedTextArray = decodedHtml.match(/(?<=<[^>]+>)([^<]+)(?=<\/[^>]+)/g) || []; console.log(extractedTextArray); // Output: ['ppshein', 'male', 'javascript']
How the Regex Works
Let’s break down the pattern /(?<=<[^>]+>)([^<]+)(?=<\/[^>]+)/g:
(?<=<[^>]+>): Positive lookbehind – makes sure we’re positioned right after an opening tag (starts with<, ends with>)([^<]+): Capture group – grabs all characters that aren’t<(this is the actual text inside the tag)(?=<\/[^>]+): Positive lookahead – ensures we’re positioned right before a closing tag (starts with</, ends with>)- The
gflag tells the regex to find all matches instead of just the first one.
Edge Case Handling
If there’s a chance some tags might be empty (like <b></b>), add a quick filter to remove empty strings from the final array:
const cleanedArray = extractedTextArray.filter(text => text.trim() !== '');
内容的提问来源于stack exchange,提问作者Pyae Phyoe Shein

