JavaScript:移除HTML标签、修改文本并重新插入标签
Great question—this is a common challenge when working with HTML as raw text (without DOM parsing) and needing to retain tag placement after modifying content. Let’s walk through a robust approach that addresses all your requirements:
Core Approach
Instead of trying to strip tags and track their absolute positions (which breaks when text is edited), we’ll split the HTML into alternating segments of plain text and HTML tags. This way, you can modify only the text segments, then rejoin everything to restore the original tag structure in the correct places. We’ll also handle escaped </> characters to avoid false positive matches.
Step 1: Temporarily Replace Escaped Entities
First, we’ll swap < and > with unique temporary placeholders. This prevents our regex from mistaking these escaped characters for actual HTML tag boundaries:
// Use unique placeholders that won't appear in your content const tempLT = '__TEMP_LT__'; const tempGT = '__TEMP_GT__'; let processedHtml = originalHtml.replace(/</g, tempLT).replace(/>/g, tempGT);
Step 2: Split HTML into Text/Tag Segments
We’ll use an improved regex to split the HTML into an array of segments, where each segment is either plain text or a full HTML tag. The regex handles tags with quoted attributes (e.g., <input type="text">) to avoid partial matches:
const segmentRegex = /(<(?:[^>"']|"[^"]*"|'[^']*')*>)/g; const segments = []; let lastIndex = 0; let match; // Iterate through all tag matches while ((match = segmentRegex.exec(processedHtml)) !== null) { // Add the plain text before the current tag if (match.index > lastIndex) { segments.push({ type: 'text', content: processedHtml.slice(lastIndex, match.index) }); } // Add the tag itself segments.push({ type: 'tag', content: match[0] }); lastIndex = match.index + match[0].length; } // Add any remaining plain text after the last tag if (lastIndex < processedHtml.length) { segments.push({ type: 'text', content: processedHtml.slice(lastIndex) }); }
Step 3: Modify the Plain Text Segments
Now you can safely edit only the text segments—tags remain untouched. For example, here’s how to convert all text to uppercase (replace this with your own editing logic):
const modifiedSegments = segments.map(segment => { if (segment.type === 'text') { // Your custom text modification goes here return { ...segment, content: segment.content.toUpperCase() }; } // Leave tags unchanged return segment; });
Step 4: Restore Escaped Entities and Reconstruct HTML
Finally, swap the temporary placeholders back to </> and join all segments to get your modified HTML with tags in their original positions:
const finalHtml = modifiedSegments .map(seg => seg.content) .join('') .replace(new RegExp(tempLT, 'g'), '<') .replace(new RegExp(tempGT, 'g'), '>');
Why This Works
- Avoids DOMParser issues: Pure string manipulation means no reliance on browser DOM APIs, making it perfect for external site scripts (like user scripts).
- No false tag matches: The temporary placeholder swap ensures escaped
</>characters aren’t mistaken for real tags. - Perfect tag position retention: By keeping tags and text in a sequential array, modifying text doesn’t disrupt where tags should be inserted—they stay in their original relative positions.
Edge Cases Handled
- Tags with quoted attributes (e.g.,
<a href="https://example.com?x=y&z=1">) - Multiple consecutive tags
- Text with mixed escaped entities and real tags
- Empty text segments (e.g., tags at the start/end of HTML)
内容的提问来源于stack exchange,提问作者Magnus

