You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JavaScript:移除HTML标签、修改文本并重新插入标签

Solution: Preserve HTML Tag Positions While Editing Text

Great question—this is a common challenge when working with HTML as raw text (without DOM parsing) and needing to retain tag placement after modifying content. Let’s walk through a robust approach that addresses all your requirements:

Core Approach

Instead of trying to strip tags and track their absolute positions (which breaks when text is edited), we’ll split the HTML into alternating segments of plain text and HTML tags. This way, you can modify only the text segments, then rejoin everything to restore the original tag structure in the correct places. We’ll also handle escaped </> characters to avoid false positive matches.

Step 1: Temporarily Replace Escaped Entities

First, we’ll swap &lt; and &gt; with unique temporary placeholders. This prevents our regex from mistaking these escaped characters for actual HTML tag boundaries:

// Use unique placeholders that won't appear in your content
const tempLT = '__TEMP_LT__';
const tempGT = '__TEMP_GT__';
let processedHtml = originalHtml.replace(/&lt;/g, tempLT).replace(/&gt;/g, tempGT);

Step 2: Split HTML into Text/Tag Segments

We’ll use an improved regex to split the HTML into an array of segments, where each segment is either plain text or a full HTML tag. The regex handles tags with quoted attributes (e.g., <input type="text">) to avoid partial matches:

const segmentRegex = /(<(?:[^>"']|"[^"]*"|'[^']*')*>)/g;
const segments = [];
let lastIndex = 0;
let match;

// Iterate through all tag matches
while ((match = segmentRegex.exec(processedHtml)) !== null) {
  // Add the plain text before the current tag
  if (match.index > lastIndex) {
    segments.push({
      type: 'text',
      content: processedHtml.slice(lastIndex, match.index)
    });
  }
  // Add the tag itself
  segments.push({
    type: 'tag',
    content: match[0]
  });
  lastIndex = match.index + match[0].length;
}

// Add any remaining plain text after the last tag
if (lastIndex < processedHtml.length) {
  segments.push({
    type: 'text',
    content: processedHtml.slice(lastIndex)
  });
}

Step 3: Modify the Plain Text Segments

Now you can safely edit only the text segments—tags remain untouched. For example, here’s how to convert all text to uppercase (replace this with your own editing logic):

const modifiedSegments = segments.map(segment => {
  if (segment.type === 'text') {
    // Your custom text modification goes here
    return { ...segment, content: segment.content.toUpperCase() };
  }
  // Leave tags unchanged
  return segment;
});

Step 4: Restore Escaped Entities and Reconstruct HTML

Finally, swap the temporary placeholders back to &lt;/&gt; and join all segments to get your modified HTML with tags in their original positions:

const finalHtml = modifiedSegments
  .map(seg => seg.content)
  .join('')
  .replace(new RegExp(tempLT, 'g'), '&lt;')
  .replace(new RegExp(tempGT, 'g'), '&gt;');

Why This Works

  • Avoids DOMParser issues: Pure string manipulation means no reliance on browser DOM APIs, making it perfect for external site scripts (like user scripts).
  • No false tag matches: The temporary placeholder swap ensures escaped </> characters aren’t mistaken for real tags.
  • Perfect tag position retention: By keeping tags and text in a sequential array, modifying text doesn’t disrupt where tags should be inserted—they stay in their original relative positions.

Edge Cases Handled

  • Tags with quoted attributes (e.g., <a href="https://example.com?x=y&z=1">)
  • Multiple consecutive tags
  • Text with mixed escaped entities and real tags
  • Empty text segments (e.g., tags at the start/end of HTML)

内容的提问来源于stack exchange,提问作者Magnus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:05:24