You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP DomXpath如何提取首行嵌套表格的非标准HTML表格数据

使用DomXpath提取该非标准嵌套表格的实现方案

该页面表头未使用标准rowspan/colspan语义化属性,而是在首行嵌入子表格实现双层表头,可按「提取合并表头-映射组装正文」两步实现需求,代码兼容后续表格列数调整,不需要硬编码表头文本。

初始化DOM与Xpath实例

首先加载页面HTML,处理编码避免葡语特殊字符乱码,同时抑制非标准HTML结构的加载警告:

// 拉取目标页面HTML内容,可选用curl、file_get_contents等方式获取
$html = file_get_contents('目标页面地址');
$dom = new DOMDocument();
// 转码避免特殊字符乱码
$html = mb_convert_encoding($html, 'HTML-ENTITIES', 'UTF-8');
@$dom->loadHTML($html);
$xpath = new DOMXPath($dom);

动态提取合并双层表头

逻辑说明:

  • 先提取主表格首行中,嵌套子表格之前的独立列(即Data、DataFim两列)
  • 解析嵌套子表格的两行内容:第一行为大分组名,按colspan属性展开对应列数;第二行为子列名
  • 将分组名和子列名按分组名 - 子列名的格式拼接,和前面的独立列合并得到最终8列表头
// 定位主数据表格
$mainTable = $xpath->query('//table[not(ancestor::tr)]')->item(0);
$firstHeaderRow = $xpath->query('./tr[1]', $mainTable)->item(0);

$topIndependentColumns = [];
$nestedHeaderTable = null;
// 遍历首行单元格,拆分独立列和嵌套子表格
foreach ($xpath->query('./td', $firstHeaderRow) as $td) {
    $nestedTable = $xpath->query('.//table', $td);
    if ($nestedTable->length > 0) {
        $nestedHeaderTable = $nestedTable->item(0);
        break;
    }
    $cellText = trim(preg_replace('/\s+/', ' ', $td->nodeValue));
    if (!empty($cellText)) {
        $topIndependentColumns[] = $cellText;
    }
}

// 处理嵌套子表格的双层表头
$nestedGroupRow = $xpath->query('./tr[1]/td', $nestedHeaderTable);
$nestedSubColRow = $xpath->query('./tr[2]/td', $nestedHeaderTable);

$expandedGroups = [];
// 按colspan展开分组名,匹配对应子列数量
foreach ($nestedGroupRow as $groupTd) {
    $colspan = (int)$groupTd->getAttribute('colspan') ?: 1;
    $groupName = trim(preg_replace('/\s+/', ' ', $groupTd->nodeValue));
    for ($i = 0; $i < $colspan; $i++) {
        $expandedGroups[] = $groupName;
    }
}

// 拼接分组名和子列名
$mergedNestedColumns = [];
foreach ($nestedSubColRow as $idx => $subColTd) {
    $subColText = trim(preg_replace('/\s+/', ' ', $subColTd->nodeValue));
    $mergedNestedColumns[] = $expandedGroups[$idx] . ' - ' . $subColText;
}

// 合并得到最终表头数组,共8项
$columns_extracted_result = array_merge($topIndependentColumns, $mergedNestedColumns);

组装正文关联数组

跳过主表格首行(表头行),遍历剩余数据行,按表头索引映射单元格值,自动跳过空行、分隔行:

$table = [];
// 取主表格中除首行表头外的所有行
$dataRows = $xpath->query('./tr[position() > 1]', $mainTable);
foreach ($dataRows as $row) {
    $cells = $xpath->query('./td', $row);
    // 单元格数量和表头长度不匹配则跳过(空行、分隔行)
    if ($cells->length !== count($columns_extracted_result)) {
        continue;
    }
    $rowItem = [];
    foreach ($cells as $colIdx => $cell) {
        $cellValue = trim(preg_replace('/\s+/', ' ', $cell->nodeValue));
        $rowItem[$columns_extracted_result[$colIdx]] = $cellValue;
    }
    $table[] = $rowItem;
}

适配说明

  • 代码自动识别单元格colspan属性,若后续页面调整储蓄规则分组、增减列数,不需要修改提取逻辑即可正常生成表头
  • 所有文本提取时都做了空白字符清理,避免换行、连续空格导致键名、数据出现多余不可见字符
  • 若页面编码调整,只需修改mb_convert_encoding的源编码参数即可解决乱码问题

内容的提问来源于stack exchange,提问作者celsowm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 05:09:20