You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP技术问询:如何从$data变量的HTML数据提取产品描述与成分?

PHP提取HTML中指定区块内容的实现方案

当然有可行的办法!在PHP里处理HTML内容提取,最稳妥的方式是用DOM扩展(别用正则,复杂HTML下正则很容易翻车)。下面我给你一套具体的实现思路和代码,完美适配你提到的需求:

核心思路

我们利用DOMDocument加载HTML内容,再通过DOMXPath精准定位到「Product Description」和「Ingredients」对应的区块,然后提取它们后面的内容。这种方法能应对大多数不规范的HTML结构,比正则靠谱得多。

具体代码实现

假设你的$data变量里的HTML结构类似这样(如果实际结构不同,只需微调XPath规则即可):

<div class="product-details">
  <h3>Product Description</h3>
  <p>It can help improve brain and memory function</p>
  <p>Its good for health</p>
  <h3>Ingredients</h3>
  <p>Lecithin (product of soya), Gelatin, Glycerine. ...</p>
</div>

对应的PHP代码:

<?php
// 你的HTML数据
$data = '<div class="product-details">
  <h3>Product Description</h3>
  <p>It can help improve brain and memory function</p>
  <p>Its good for health</p>
  <h3>Ingredients</h3>
  <p>Lecithin (product of soya), Gelatin, Glycerine. ...</p>
</div>';

// 初始化DOM文档,忽略HTML解析错误(很多网页HTML不规范)
libxml_use_internal_errors(true);
$dom = new DOMDocument();
$dom->loadHTML($data);
libxml_clear_errors();

// 创建XPath对象用于节点查询
$xpath = new DOMXPath($dom);

// 提取产品描述($pdesc)
$pdesc = '';
// 先定位到文本为"Product Description"的标题节点(这里假设是h3,实际根据你的HTML调整)
$descTitle = $xpath->query('//h3[text()="Product Description"]')->item(0);
if ($descTitle) {
    // 遍历标题后面的兄弟节点,直到遇到下一个标题或无节点
    $nextNode = $descTitle->nextSibling;
    while ($nextNode) {
        // 只提取p标签的内容(根据实际结构调整标签类型)
        if ($nextNode->nodeType === XML_ELEMENT_NODE && $nextNode->tagName === 'p') {
            $pdesc .= $dom->saveHTML($nextNode) . ' ';
        }
        // 遇到下一个h3标题就停止遍历
        if ($nextNode->nodeType === XML_ELEMENT_NODE && $nextNode->tagName === 'h3') {
            break;
        }
        $nextNode = $nextNode->nextSibling;
    }
    $pdesc = trim($pdesc);
}

// 提取成分($ingredient)
$ingredient = '';
$ingTitle = $xpath->query('//h3[text()="Ingredients"]')->item(0);
if ($ingTitle) {
    $nextNode = $ingTitle->nextSibling;
    while ($nextNode) {
        if ($nextNode->nodeType === XML_ELEMENT_NODE && $nextNode->tagName === 'p') {
            $ingredient .= $dom->saveHTML($nextNode) . ' ';
        }
        if ($nextNode->nodeType === XML_ELEMENT_NODE && $nextNode->tagName === 'h3') {
            break;
        }
        $nextNode = $nextNode->nextSibling;
    }
    $ingredient = trim($ingredient);
}

// 测试输出结果
var_dump($pdesc);
var_dump($ingredient);
?>

灵活调整说明

如果你的HTML结构和示例不同,只需修改XPath查询规则即可:

  • 若标题不是h3,而是h2或带class的div:比如//h2[@class="section-header" and text()="Product Description"]
  • 若内容被包裹在容器div里:比如找到标题后,取它的下一个兄弟div容器,再提取里面的p标签:$xpath->query('//h3[text()="Product Description"]/following-sibling::div[1]//p')

这种方法的优势是能稳定处理各种HTML结构变化,不像正则表达式容易因为换行、额外属性就匹配失败。

内容的提问来源于stack exchange,提问作者Rand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:04:02