PHP DOMDocument loadHTML丢失WordPress文章<br/>换行的解决方法
问题:DOMDocument处理WordPress内容时
<br>标签换行丢失 使用PHP DOMDocument的loadHTML方法处理WordPress文章内容时,发现<br />标签生成的换行被合并,所有文本输出为一行,可读性极差。
期望输出:
Specifications: Name: VR glasses Type: virtual reality glasses Model: for VRGPRO+ Lens type: blue coated lens
实际输出:
Specifications:Name: VR glassesType: virtual reality glassesModel: for VRGPRO+Lens type: blue coated lens
已尝试以下方法,均未解决问题:
- 移除
LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD | LIBXML_PARSEHUGE | LIBXML_COMPACT | LIBXML_NOBLANKS参数 - 在
strip_tags中允许<br>标签 - 简化处理代码
主题中的相关代码:
$dom = new DOMDocument; $dom->loadHTML(strip_tags(mb_convert_encoding(get_the_content(), 'HTML-ENTITIES', 'UTF-8'), '<img>,<div>,<table>,<thead>,<tbody>,<tfoot>,<tr>,<th>,<td>,<ul>,<ol>,<li>,<strong>,<em>,<h3>,<h4>,<h5>,<h6>,<p>'), LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD | LIBXML_PARSEHUGE | LIBXML_COMPACT | LIBXML_NOBLANKS); $xpath = new DOMXPath($dom); foreach ($xpath->query('//@*[not(name()="src")]') as $attr) { $attr->parentNode->removeAttribute($attr->nodeName); } $images = $dom->getElementsByTagName('img'); foreach ( $images as $image ) { $properlink = rtrim(preg_replace ('//','https:',$image->getAttribute('src'),1),'/'); $image->setAttribute( 'alt', esc_attr( wp_trim_words( get_the_title(), 10,'') ) ); $image->setAttribute( 'data-src', $properlink ); $image->removeAttribute('src'); $image->setAttribute( 'style', 'width:100%;max-width:100%;' ); } while (($node_list = $xpath->query('//*[not(*) and not(@*) and not(text()[normalize-space()])]')) && $node_list->length) { foreach ($node_list as $node) { $node->parentNode->removeChild($node); } } echo $dom->saveHTML();
更新信息:
- WordPress文章中确实存在
<br />标签,格式如下:
Specifications:<br />Name: VR glasses<br />Type: virtual reality glasses<br />Model: for VRGPRO+<br />Lens type: blue coated lens<br />
尝试在
strip_tags中允许<br>或<br/>标签,要么页面空白,要么换行仍丢失。简化代码后问题依旧:
$content = get_the_content(); $dom = new DOMDocument(); $dom->loadHTML( strip_tags($content, '<img><br><div><table><thead><tbody><tfoot><tr><th><td><ul><ol><li><strong><em><h3><h4><h5><h6><p>'), LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD | LIBXML_PARSEHUGE | LIBXML_COMPACT | LIBXML_NOBLANKS ); $html = $dom->saveHTML(); $html = str_replace('<br>', "<br>\n", $html); // if you want newlines echo $html;
- 已提供典型的WordPress文章内容示例。
内容的提问来源于stack exchange,提问作者Johnny Bravo
相关产品推荐
相关产品推荐

