如何用XPath提取表格中除首个<tr>标签外的<tr>及<td>内容?
正确提取表格内容的XPath方案
需求背景
需提取目标HTML表格中除第一个<tr>标签外的所有<tr>及<td>内容,当前使用的/table/tbody/tr/td会包含表头行,需调整XPath表达式。
原HTML表格代码
<table border="1px" cellpadding="10px" style="background-color:rgb(247, 247, 247); border-collapse:collapse; border:1px; height:450px; max-width:650px; width:100%"> <tbody> <tr> <td colspan="2" style="background-color:rgb(255, 255, 255); border-color:rgb(236, 238, 239); height:50px; width:628px"> <span style="color:#000000"><span style="font-size:14px"><span style="font-family:trebuchet ms,helvetica,sans-serif"><strong>Descrizione tecnica</strong></span></span></span> </td> </tr> <tr> <td style="background-color:rgb(244, 249, 252); border-color:rgb(236, 238, 239); height:50px; width:304px"> <span style="color:#000000"><span style="font-size:14px"><span style="font-family:trebuchet ms,helvetica,sans-serif"><strong>CPU Dissipatore</strong></span></span></span> </td> <td style="background-color:rgb(244, 249, 252); border-color:rgb(236, 238, 239); height:50px; width:303px"> <span style="color:#000000"><span style="font-size:14px"><span style="font-family:trebuchet ms,helvetica,sans-serif">Intel i7-11700K<p></p> Dissipatore a Liquido 240mm </span></span> </span> </td> </tr> </tbody> </table>
修正后的XPath方案
方案1:通过位置索引过滤(通用场景)
适用于表格有多个数据行,需要跳过第一个表头行的情况:
/table/tbody/tr[position() > 1]/td
如果确定只有1个数据行(如示例),也可以直接指定索引:
/table/tbody/tr[2]/td
方案2:通过表头行的属性特征过滤
利用第一个tr内的td带有colspan="2"的特征,精准排除表头行:
/table/tbody/tr[td/@colspan != 2]/td
说明
- 方案1通过
position() > 1选中tbody下所有位置大于1的tr,直接跳过第一个表头行,提取其下所有td内容,通用性强。 - 方案2通过属性特征筛选,适合表头行有明显标识(如colspan、特定class等)的场景,准确性更高。
内容的提问来源于stack exchange,提问作者Linux
相关产品推荐
相关产品推荐

