PHP strip_tags无法保留DOM中<img>标签?WordPress内容处理问题排查
WordPress摘要处理中图片丢失的问题排查与修复
问题重现
为了让WordPress博客摘要仅保留<img>及指定HTML元素,我替换了the_content()改用DOM处理方案,但其他元素正常显示,唯独图片完全消失。
原始处理代码
// the_content(); $dom = new DOMDocument; @$dom->loadHTML(strip_tags(mb_convert_encoding(get_the_content(), 'HTML-ENTITIES', 'UTF-8'), '<img>|<p>|<div>|<table>|<thead>|<tbody>|<tfoot>|<tr>|<th>|<td>|<ul>|<ol>|<li>|<strong>|<em>|<h3>|<h4>|<h5>|<h6>|<b>|<i>|<span>'), LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD); $xpath = new DOMXPath($dom); //removes all attributes $nodes = $xpath->query('//@*'); foreach ($nodes as $node) { $node->parentNode->removeAttribute($node->nodeName); } //removes <p> </p> $nodeList = $xpath->query("//p[normalize-space(.)=\"\xC2\xA0\"]"); # foreach($nodeList as $node) { $node->parentNode->removeChild($node); } //remove all elements attributes except for img > src foreach ($xpath->query('//@*[not(name()="src")]') as $attr) { //@*[not(name()="src" or name()="href")] $attr->parentNode->removeAttribute($attr->nodeName); } //remove empty html tags while (($node_list = $xpath->query('//*[not(*) and not(@*) and not(text()[normalize-space()])]')) && $node_list->length) { foreach ($node_list as $node) { $node->parentNode->removeChild($node); } } echo wpautop($dom->saveHTML(), true);
原始文章内容(含图片)
<p> <span style="font-size:22px"> <strong>power supply:Transmitter: 23A 12V, the battery can be used for one year (calculated by 20 times/1 day), the buyer needs to configure it by himself</strong> </span> </p> <p> <span style="font-size:22px"> <strong>Receiver: 2*1.5V (AA battery), need to be equipped by the buyer</strong> </span> </p> <p> <br/> </p> <p> <img src="//ae01.alicdn.com/kf/S3f9212e14ba742fda4daffb1af2d4607v.jpg"/> <img src="//ae01.alicdn.com/kf/S9dee11d0fc6a4b1f916c3c8c2db8a63e9.jpg"/> </p>
实际输出(图片消失)
<p> <span> <strong>power supply:Transmitter: 23A 12V, the battery can be used for one year (calculated by 20 times/1 day), the buyer needs to configure it by himself</strong> </span> </p> <p> <span> <strong>Receiver: 2*1.5V (AA battery), need to be equipped by the buyer</strong> </span> </p>
问题根源
代码执行顺序和逻辑错误导致图片被误删:
- 第一步全删所有属性:先执行的
//removes all attributes代码,把所有元素的属性(包括<img>的src)全部删除,此时<img>变成无属性的空标签<img>。 - 保留src的逻辑无效:后续试图保留img的src属性,但此时src已经被删掉,这段代码找不到目标属性,完全不起作用。
- 空标签删除逻辑清除img:最后一段删除空标签的逻辑,会移除「没有子元素、没有属性、没有有效文本」的元素,无属性的
<img>恰好符合条件,被彻底删除。
修复方案
调整代码逻辑,先精准保留img的src属性,再删除其他不需要的属性,同时删除多余的「全删属性」步骤:
修改后的完整代码
// 替换原the_content() $dom = new DOMDocument; @$dom->loadHTML(strip_tags(mb_convert_encoding(get_the_content(), 'HTML-ENTITIES', 'UTF-8'), '<img>|<p>|<div>|<table>|<thead>|<tbody>|<tfoot>|<tr>|<th>|<td>|<ul>|<ol>|<li>|<strong>|<em>|<h3>|<h4>|<h5>|<h6>|<b>|<i>|<span>'), LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD); $xpath = new DOMXPath($dom); // 移除<p> </p>这类空段落 $nodeList = $xpath->query("//p[normalize-space(.)=\"\xC2\xA0\"]"); # 匹配 foreach($nodeList as $node) { $node->parentNode->removeChild($node); } // 精准删除不需要的属性: // 1. 非img元素的所有属性 // 2. img元素的非src属性 $attrsToRemove = $xpath->query('//*[not(self::img)]/@* | //img/@*[not(name()="src")]'); foreach ($attrsToRemove as $attr) { $attr->parentNode->removeAttribute($attr->nodeName); } // 删除空标签(此时img有src属性,不会被删除) while (($node_list = $xpath->query('//*[not(*) and not(@*) and not(text()[normalize-space()])]')) && $node_list->length) { foreach ($node_list as $node) { $node->parentNode->removeChild($node); } } echo wpautop($dom->saveHTML(), true);
修复效果
修改后的代码会保留带有src属性的<img>标签,同时正常处理其他指定元素,输出结果会包含所有图片,且其他元素的多余属性被清除。
内容的提问来源于stack exchange,提问作者Johnny Bravo
相关产品推荐
相关产品推荐

