TinyMCE自动编码HTML标签:如何禁用编码以原样保存数据?
我之前也踩过TinyMCE实体编码的坑,结合你的配置和问题描述,大概率是几个配置项冲突或者粘贴处理逻辑导致的双重转义,给你几个针对性的解决方案:
1. 先排查粘贴处理的逻辑
你在paste_preprocess里调用了html_decode,这很可能是问题的源头。假设你输入的内容是<pre><code><b>Test</b></code></pre>,html_decode会把<和>解码成<和>,变成<pre><code><b>Test</b></code></pre>。之后TinyMCE即使设置了entity_encoding: "raw",在保存时可能还是会对代码块内的<>进行转义,最终就变成了&lt;b&gt;test&lt;/b&gt;这种双重编码的结果。
建议先注释掉html_decode的调用,调整后的paste_preprocess逻辑:
paste_preprocess: function(p1, precontent){ var clean_content = clear_content(precontent.content); // 去掉html_decode,直接使用清理后的内容 precontent.content = clean_content; }
(如果clear_content函数本身包含转义/解码逻辑,也需要检查它是否会干扰原始内容)
2. 强化实体编码的配置
仅设置entity_encoding: "raw"可能不够,还需要禁用TinyMCE默认的实体转义列表,同时调整粘贴插件的行为,避免自动处理内容:
tinymce.init({ selector: '#post-message', mode: "specific_textareas", height: 500, menubar: false, plugins: 'paste print preview searchreplace autolink directionality visualblocks visualchars fullscreen image link media template codesample table charmap hr pagebreak nonbreaking anchor toc insertdatetime advlist lists textcolor wordcount imagetools contextmenu colorpicker textpattern help', theme_advanced_buttons3_add : "pastetext,pasteword,selectall", toolbar: 'bold italic link | numlist bullist', paste_word_valid_elements: "b,i,p,a[href],ol,ul,li,em,br", // 关键配置:禁用实体转义 entity_encoding: "raw", entities: "", // 清空默认的实体转义列表,确保不转义任何实体 // 调整粘贴插件,避免自动修改内容 paste_auto_cleanup_on_paste: false, paste_remove_styles: false, paste_remove_styles_if_webkit: false, paste_strip_class_attributes: 'none', paste_preprocess: function(p1, precontent){ var clean_content = clear_content(precontent.content); precontent.content = clean_content; }, branding: false });
3. 确保获取内容的方式正确
当你从编辑器中获取内容时,要明确指定raw格式,避免TinyMCE在返回内容时额外处理:
// 获取原始内容,不做任何编码处理 var editorContent = tinymce.get('post-message').getContent({ format: 'raw' });
4. 最后检查后端是否有额外转义
有时候问题不在前端,后端框架(比如PHP的htmlspecialchars、Python的escape等)会自动对存入数据库的内容进行转义,导致双重编码。可以先把前端获取到的内容打印出来,确认是否已经是正确的原始内容,再检查后端保存逻辑。
内容的提问来源于stack exchange,提问作者MrCujo

