You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用PHP XPath提取Joomla/Hikashop描述中的Tab标签及对应内容

实现方案

你可以基于正在使用的DomCrawler组件,通过DOM节点顺序遍历的方式实现精准提取,完全避开正则匹配的容错问题:

// 先匹配所有tab标记节点和结束标记节点,normalize-space自动忽略文本前后空白
$allMarkers = $crawler->filterXpath('//p[
    contains(normalize-space(text()), "{tab=") 
    or contains(normalize-space(text()), "{/tabs}")
]');

$tabs = [];
$prevMarker = null;

$allMarkers->each(function ($node) use (&$tabs, &$prevMarker) {
    $text = trim($node->text());
    // 处理tab起始标记,提取tab名称
    if (substr($text, 0, 5) === '{tab=') {
        $tabName = rtrim(substr($text, 5), '}');
        $tabs[] = [
            'name' => $tabName,
            'content' => ''
        ];
        $prevMarker = $node;
        return;
    }
    // 处理结束标记,收集最后一个tab的内容
    if ($text === '{/tabs}' && $prevMarker !== null) {
        $lastTab = &$tabs[array_key_last($tabs)];
        $content = '';
        $nextNode = $prevMarker->nextSibling();
        while ($nextNode !== null && !$nextNode->equals($node)) {
            $content .= $nextNode->outerHtml();
            $nextNode = $nextNode->nextSibling();
        }
        $lastTab['content'] = $content;
        $prevMarker = null;
        return;
    }
    // 处理两个tab标记之间的内容
    if ($prevMarker !== null) {
        $currentTab = &$tabs[array_key_last($tabs)];
        $content = '';
        $nextNode = $prevMarker->nextSibling();
        while ($nextNode !== null && !$nextNode->equals($node)) {
            // 若需要过滤空p标签,可在此加判断:if(trim($nextNode->text()) !== '') 再拼接
            $content .= $nextNode->outerHtml();
            $nextNode = $nextNode->nextSibling();
        }
        $currentTab['content'] = $content;
        $prevMarker = $node;
    }
});

// 输出结果示例
foreach ($tabs as $tab) {
    echo "Tab名称:{$tab['name']}\n";
    echo "内容:{$tab['content']}\n\n";
}

关键说明

  1. 完全基于DOM树结构遍历,不管内容是p、ul、div还是任意嵌套标签,都能完整保留原有结构提取,不会出现正则匹配的漏配、错配问题
  2. Xpath用normalize-space()处理文本匹配,自动忽略文本前后的空格、换行、缩进,适配不同排版的导出数据
  3. 代码兼容PHP7.4+,如果使用原生DOMDocument实现,逻辑完全一致,仅需将对应API替换为原生DOM的nextSibling、C14N()(节点HTML输出)即可
  4. 如需过滤内容里的空标签、冗余样式,在拼接content时增加对应判断逻辑即可

内容的提问来源于stack exchange,提问作者Darren

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 19:24:04