PHP DOMDocument解析script时丢失部分代码的问题咨询
DOMDocument解析script内容丢失的解决办法
当使用PHP的DOMDocument解析包含嵌套拆分script标签的HTML时,比如:
$html1 = "<script>document.write('<scr'+'ipt>alert(123);</scr'+'ipt>')</script>"; $dom = new DOMDocument('1.0', 'utf-8'); $dom->loadHTML($html1); $html2 = $dom->saveHTML();
得到的结果中会丢失</scr'+'ipt>部分,输出变为:
<html><head><script>document.write('\<scr'+'ipt\>alert(123);')</script></head></html>
问题原因
DOMDocument默认使用HTML解析器,该解析器会将</script>作为script标签的结束标记。即使这个标记是拆分在字符串中的(如</scr'+'ipt>),解析器也会误将</scr识别为结束标记的起始,导致后续内容被截断丢弃。
解决方案
方法一:使用XML解析器+CDATA包裹
将script内容放在CDATA块中,使用loadXML替代loadHTML解析,XML解析器会将CDATA内的内容视为纯文本,不会进行标签匹配:
$html1 = "<script><![CDATA[document.write('<scr'+'ipt>alert(123);</scr'+'ipt>')]]></script>"; $dom = new DOMDocument('1.0', 'utf-8'); libxml_disable_entity_loader(true); // 禁止外部实体加载,提升安全性 $dom->loadXML($html1); $html2 = $dom->saveXML(); // 移除CDATA标签,还原为HTML格式的script内容 $html2 = str_replace(['<![CDATA[', ']]>'], '', $html2); // 输出结果 echo htmlspecialchars($html2);
执行后会保留完整的script内容,输出为:
<script>document.write('<scr'+'ipt>alert(123);</scr'+'ipt>')</script>
方法二:临时编码script内容
解析前将script内容编码为base64,避免解析器误识别,解析完成后再解码还原:
$html1 = "<script>document.write('<scr'+'ipt>alert(123);</scr'+'ipt>')</script>"; // 提取script内容并编码 preg_replace_callback('/<script>(.*?)<\/script>/s', function($matches) use (&$html1) { $encoded = base64_encode($matches[1]); $html1 = str_replace($matches[0], "<script>__ENCODED__{$encoded}__ENCODED__</script>", $html1); }, $html1); $dom = new DOMDocument('1.0', 'utf-8'); libxml_use_internal_errors(true); // 忽略HTML解析过程中的非致命错误 $dom->loadHTML($html1); $html2 = $dom->saveHTML(); // 解码还原script内容 $html2 = preg_replace_callback('/<script>__ENCODED__(.*?)__ENCODED__<\/script>/s', function($matches) { return "<script>" . base64_decode($matches[1]) . "</script>"; }, $html2); echo htmlspecialchars($html2);
这种方法不需要修改原始HTML的结构,适合处理复杂的HTML内容。
内容的提问来源于stack exchange,提问作者RomanG
相关产品推荐
相关产品推荐

