You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP使用XPath解析嵌套HTML标签提取字段生成指定结构数组

解决方法

你当前已经正确获取到了所有class为tile的article节点,只需要在每个节点的上下文范围内单独查询对应子字段、做格式转换即可。

完整可运行代码

$dom = new DOMDocument();
// 抑制非规范HTML带来的解析警告
@$dom->loadHTML($html['html']);
$xpath = new DOMXPath($dom);
$parsedArray = [];

$nodelist = $xpath->query("//article[contains(@class, 'tile')]");

foreach ($nodelist as $node) {
    // 提取原始日期并转换格式
    $dateNode = $xpath->query(".//label[contains(@class, 'label-date--blue')]", $node)->item(0);
    $rawDate = trim($dateNode->nodeValue);
    $date = DateTime::createFromFormat('d.m.Y', $rawDate)->format('Y-m-d');

    // 提取标题和对应链接
    $titleANode = $xpath->query(".//h4/a[contains(@class, 'link-color-black')]", $node)->item(0);
    $title = trim($titleANode->nodeValue);
    $link = trim($titleANode->getAttribute('href'));

    // 提取内容文本
    $contentNode = $xpath->query(".//p[contains(@class, 'tile-content__paragraph--gray')]", $node)->item(0);
    $content = trim($contentNode->nodeValue);

    // 组装进结果数组
    $parsedArray[] = [
        'title' => $title,
        'link' => $link,
        'date' => $date,
        'content' => $content
    ];
}

// 输出结果验证
echo '<pre>';
var_dump($parsedArray);
echo '</pre>';

关键逻辑说明

  • 子节点XPath查询前缀加.,限定查询范围为当前遍历的单个article节点,不会匹配其他article下的同类型字段
  • 使用getAttribute('属性名')方法提取HTML标签的属性值,比如链接的href属性
  • 字段提取后统一加trim()处理,清除文本前后多余的空格、换行符
  • 日期格式转换使用DateTime::createFromFormat精准匹配原日.月.年的格式,再输出目标年-月-日格式,避免解析错误

兼容优化(可选)

如果实际运行场景中可能存在部分字段缺失的情况,可以在提取节点前增加非空判断,避免抛出未定义错误:

$dateNode = $xpath->query(".//label[contains(@class, 'label-date--blue')]", $node)->item(0);
$rawDate = $dateNode ? trim($dateNode->nodeValue) : '';
$date = $rawDate ? DateTime::createFromFormat('d.m.Y', $rawDate)->format('Y-m-d') : '';

内容的提问来源于stack exchange,提问作者Igor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 11:06:04