You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用sanitize-html npm包清理HTML字符串无效,求助排查问题

sanitize-html 清理HTML无效,危险属性未被移除的解决办法

问题重现

使用sanitize-html处理HTML字符串时,输入包含onerror恶意事件的图片标签,结果危险属性未被过滤,还自动添加了<p>标签。

代码片段

if (noteContent) {
  const resultContent = sanitize(noteContent);
  console.log(resultContent);
} else {
  setErrorMessages((prevState) => ["note content cannot be empty"]);
}

输入输出示例

  • 输入:<img src=? onerror="alert('hello')" />
  • 输出:<p><img src=? onerror="alert('hello')" /></p>

原因分析

  1. 输入为转义后的HTML字符串:如果传入的是&lt;img...&gt;这类转义文本,sanitize-html会将其视为普通文本,不会解析为HTML标签,自然不会过滤属性。
  2. 未显式配置过滤规则:默认配置可能因版本差异或规则宽松,导致危险属性未被拦截。

解决步骤

1. 还原转义的HTML(如果输入是转义字符串)

如果你的输入是经过HTML转义的文本,先将其还原为原始HTML再处理:

// 解码转义的HTML字符
const decodedContent = noteContent.replace(/&lt;/g, '<').replace(/&gt;/g, '>');
const resultContent = sanitize(decodedContent);

2. 显式配置严格的过滤规则

手动指定允许的标签和属性,确保危险属性被移除:

const sanitizeOptions = {
  allowedTags: ['p', 'img'], // 仅允许p和img标签
  allowedAttributes: {
    img: ['src', 'alt'] // 仅允许img的src、alt属性
  }
};

const resultContent = sanitize(noteContent, sanitizeOptions);

3. 更新sanitize-html到最新版本

旧版本可能存在规则漏洞,执行升级:

npm update sanitize-html

验证效果

配置完成后,输入<img src=? onerror="alert('hello')" />会输出:

<p><img src=? /></p>

onerror属性已被成功过滤。

内容的提问来源于stack exchange,提问作者nika bedenashvili

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 08:12:42