如何用XPath获取标题后的首个<p>标签以生成FAQPage Schema
修复FAQPage结构化数据中获取对应答案的问题
问题原因
你之前尝试的following-sibling::p[1] XPath查询无效,是因为没有基于当前标题节点作为上下文执行查询,导致全局匹配到了最后一个p标签。
修复步骤
- 遍历每个标题节点时,以当前节点为上下文执行相对路径的XPath查询,获取后续第一个p标签
- 对段落内容做JSON转义处理,避免出现语法错误
- 处理标题后无p标签的空值情况
修改后的完整代码
<?php $content_postid = get_the_ID(); $content_post = get_post($content_postid); $content = $content_post->post_content; $content = apply_filters('the_content', $content); $content = str_replace(']]>', ']]>', $content); libxml_use_internal_errors(true); $dom = new DOMDocument; $dom->loadHTML('<?xml encoding="utf-8" ?>' . $content); $xp = new DOMXPath($dom); $query = "//h2[contains(., '?')] | //h3[contains(., '?')]"; $nodes = $xp->query($query); $stack = []; if ($nodes && $nodes->length > 0) { $faq_count = $nodes->length; $faq_i = 1; echo ' <script type="application/ld+json"> { "@context": "https://schema.org", "@type": "FAQPage", "mainEntity": ['; foreach($nodes as $node) { // 以当前标题为上下文,查询后续第一个p标签 $answerQuery = $xp->query('following-sibling::p[1]', $node); $answerText = ''; if ($answerQuery && $answerQuery->length > 0) { $answerText = $answerQuery->item(0)->nodeValue; } // 用json_encode处理转义,避免JSON格式错误 $questionName = json_encode($node->nodeValue); $answerTextEncoded = json_encode($answerText); $answerUrl = json_encode(get_permalink() . '#' . $node->getAttribute('id')); echo '{ "@type": "Question", "name": ' . $questionName . ', "acceptedAnswer": { "@type": "Answer", "text": ' . $answerTextEncoded . ', "url": ' . $answerUrl . ' } }'; if ($faq_i != $faq_count) : echo ','; endif; $faq_i++; } echo ']} </script>'; } ?>
代码说明
$xp->query('following-sibling::p[1]', $node):通过第二个参数指定查询上下文,确保只匹配当前标题节点后的第一个p标签json_encode():自动处理引号、换行等特殊字符,避免JSON语法错误- 增加查询结果长度判断,防止标题后无p标签时抛出错误
内容的提问来源于stack exchange,提问作者Cray
相关产品推荐
相关产品推荐

