使用PHP DOMDocument解析WordPress内容无法匹配img元素的问题
问题描述
用PHP DOMDocument解析WordPress文章内容,想只打印img元素,但代码死活输出不了任何img;要是把($item->nodeName == 'img')这个条件去掉,所有元素(包括img)都能正常打印。明明DOM里有img,为啥这个条件过滤不出来?
我的代码
function myfun($post_id) { // Get the post conent $post = get_post($post_id); $body = $post->post_content; // Parse the post content with as UTF-8 // Ref: https://www.php.net/manual/en/intro.dom.php $doc = new \DOMDocument(); $doc->loadHtml("<html><head><meta charset=\"UTF-8\"><meta http-equiv=\"Content-Type\" content=\"text/html; charset=UTF-8\"></head><body>".$body."</body></html>"); // Enumerate the DOM tree $doc_root = $doc->documentElement; enum_dom($doc_root->childNodes, 0); } function enum_dom($nodes, $level) { foreach ($nodes AS $item) { if (($item->nodeType == XML_ELEMENT_NODE) && ($item->nodeName == 'img')) { print $item->nodeName . PHP_EOL; if($item->childNodes || $item->childNodes->lenth > 0) { enum_dom($item->childNodes, $level+5); } } } }
测试用的文章内容
<img class="aligncenter wp-image-43154" title="Excel Invoice Template Site Introduction" src="https://www.sample.com/blogs/wp-content/uploads/2024/04/excel-invoice-template-site-introduction.jpg" alt="Excel Invoice Template Site Introduction" width="600" height="338" />
问题原因
PHP的DOMDocument::loadHtml()是按照HTML规则解析内容的,它会把所有HTML标签名转换成大写形式。所以你代码里判断$item->nodeName == 'img'永远匹配不上,实际img元素的nodeName是IMG。
另外提一句,你代码里还有个拼写错误:$item->childNodes->lenth应该是$item->childNodes->length,不过这个不影响当前img的匹配问题。
解决方案
有三种简单的解决方式:
- 直接把判断条件改成大写:将
$item->nodeName == 'img'替换成$item->nodeName == 'IMG',直接匹配解析后的节点名。 - 统一转小写后判断:用
strtolower($item->nodeName) == 'img',彻底避免大小写问题。 - 用更高效的内置方法:不用自己递归遍历所有节点,直接用
getElementsByTagName()获取所有img元素,代码更简洁:function myfun($post_id) { $post = get_post($post_id); $body = $post->post_content; $doc = new \DOMDocument(); $doc->loadHtml("<html><head><meta charset=\"UTF-8\"></head><body>".$body."</body></html>"); // 直接获取所有img元素 $imgs = $doc->getElementsByTagName('img'); foreach ($imgs as $img) { print $img->nodeName . PHP_EOL; } }
内容的提问来源于stack exchange,提问作者alancc
相关产品推荐
相关产品推荐

