如何移除CKEditor源码中的HTML标签及HTML属性
Hey there! Let's walk through how to handle two common tasks with your CKEditor content before storing it in the database: stripping all HTML tags entirely, and removing just the HTML attributes while keeping the tags themselves.
This is useful if you don't need any formatting at all and just want the raw text content. Here's how to do it in popular backend languages:
PHP
Use the built-in strip_tags() function—it's simple and reliable for most cases:
$plainText = strip_tags($ckeditorContent); // If you want to allow specific tags (e.g., <p>, <br>), pass them as the second argument: // $plainText = strip_tags($ckeditorContent, '<p><br>');
JavaScript/Node.js
Use the DOMParser API to parse the HTML and extract text content. For Node.js, you'll need a library like jsdom since the browser's DOMParser isn't available natively:
// Browser environment const parser = new DOMParser(); const doc = parser.parseFromString(ckeditorContent, 'text/html'); const plainText = doc.body.textContent || ""; // Node.js (with jsdom installed: npm install jsdom) const { JSDOM } = require('jsdom'); const dom = new JSDOM(ckeditorContent); const plainText = dom.window.document.body.textContent || "";
Python
Use BeautifulSoup (install via pip install beautifulsoup4) for safe HTML parsing:
from bs4 import BeautifulSoup soup = BeautifulSoup(ckeditorContent, 'html.parser') plainText = soup.get_text(strip=True)
If you want to keep basic HTML structure (like <p>, <h1>) but strip out all attributes (e.g., class, style, id), here's how to approach it:
PHP
Use a regular expression to match tags and replace them with the tag name only. Note: regex isn't perfect for all edge cases, but works for most CKEditor-generated content:
// Strip all attributes from HTML tags $cleanedContent = preg_replace('/<([a-z][a-z0-9]*)\b[^>]*>/i', '<$1>', $ckeditorContent); // For self-closing tags (like <img>, <br>), adjust the regex slightly: $cleanedContent = preg_replace('/<([a-z][a-z0-9]*)\b[^>]*\/>/i', '<$1/>', $cleanedContent);
JavaScript/Node.js
Parse the HTML, iterate over all elements, and remove their attributes:
// Browser environment const parser = new DOMParser(); const doc = parser.parseFromString(ckeditorContent, 'text/html'); const elements = doc.body.querySelectorAll('*'); elements.forEach(el => { while (el.attributes.length > 0) { el.removeAttribute(el.attributes[0].name); } }); const cleanedContent = doc.body.innerHTML; // Node.js (with jsdom) const { JSDOM } = require('jsdom'); const dom = new JSDOM(ckeditorContent); const elements = dom.window.document.body.querySelectorAll('*'); elements.forEach(el => { while (el.attributes.length > 0) { el.removeAttribute(el.attributes[0].name); } }); const cleanedContent = dom.window.document.body.innerHTML;
Python
Again, BeautifulSoup makes this straightforward:
from bs4 import BeautifulSoup soup = BeautifulSoup(ckeditorContent, 'html.parser') for tag in soup.find_all(True): tag.attrs = {} # Clear all attributes from the tag cleanedContent = str(soup)
- Avoid regex for complex HTML: Regex can break on nested tags or unusual attribute values. Always prefer DOM parsing libraries (like BeautifulSoup, jsdom) when possible.
- Handle XSS risks: If users are submitting CKEditor content, make sure to sanitize it before storing—even if you're stripping tags/attributes. Libraries like
HTMLPurifier(PHP) orbleach(Python) can help with this. - Configure CKEditor upfront: You can limit what tags/attributes users can add directly in CKEditor using
config.allowedContentorconfig.disallowedContent. This reduces the amount of cleanup you need to do later!
内容的提问来源于stack exchange,提问作者 Tarun

