如何清理Summernote生成的冗余HTML标签?JS replace失效排查
Let's break down the problems with your current code and how to resolve them:
1. Strings in JavaScript Are Immutable
The replace() method doesn’t modify the original string—it returns a new string with the changes. So you need to assign the result back to your html variable (or a new one) for the changes to stick:
html = html.replace('old', 'new'); // This works html.replace('old', 'new'); // This doesn’t change the original html variable
2. Your Search Pattern Doesn’t Match the Redundant Tags
Looking at your original HTML, the empty tags you want to remove look like this:<p> <b> <br> </b> </p>
But your replace call targets <p><br></p>—which doesn’t account for the spaces and nested <b> tags. So the pattern never matches anything.
3. String Replacement Only Works Once
Even if your pattern was correct, using a string with replace() only replaces the first occurrence. To replace all instances, you need to use a regular expression with the g (global) flag.
Corrected Code to Remove Redundant Tags
Here’s a script that fixes all these issues and targets the specific empty tags in your HTML:
var html = ` <p> <b> <br> </b> </p> <p> <b>អ្នកធានា</b> </p> <p> <b>ឈ្មោះ: ……………………………</b> </p> <p> <b>អត្តសញ្ញាណប័ណ្ណលេខៈ………………............... <span style="white-space:pre"></span>.......................................... </b> </p> <p> <b> <span style="white-space:pre"></span>ហត្ថលេខានិង ស្នាមមេដៃស្តាំ <span style="white-space:pre"></span> </b> </p> <p> <b> <br> </b> </p> <p>`; // Remove empty <p> tags containing <b><br></b> (with any whitespace) html = html.replace(/<p>\s*<b>\s*<br>\s*<\/b>\s*<\/p>/g, ''); // Remove any remaining empty <p> tags with just <br> or whitespace html = html.replace(/<p>\s*<br>\s*<\/p>/g, ''); html = html.replace(/<p>\s*<\/p>/g, ''); // Remove all instances html = html.replace(/&nbsp;/g, ''); console.log(html);
What Each Regex Does:
\s*: Matches zero or more whitespace characters (spaces, tabs, newlines) to account for inconsistent spacing in the generated HTML./g: Ensures all occurrences are replaced, not just the first.- The first regex specifically targets the
<p><b><br></b></p>pattern from your Summernote output. - The additional regexes clean up other common empty tag variations.
Bonus: More Robust Cleanup (Optional)
If you want to handle even more edge cases (like nested empty tags or varying whitespace), you can run the replacements multiple times to catch any tags that might be left after the first pass:
// Run the replacements a few times to handle nested or leftover empty tags for (let i = 0; i < 2; i++) { html = html.replace(/<p>\s*<b>\s*<br>\s*<\/b>\s*<\/p>/g, ''); html = html.replace(/<p>\s*(<br>|\s)*<\/p>/g, ''); }
This ensures any tags that might have been partially cleaned up in the first pass are fully removed.
内容的提问来源于stack exchange,提问作者Houy Narun

