You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用XPath 1.0获取最深层的<td>元素(适配lxml库)

Selecting Deepest Elements (No Descendant ) with XPath 1.0

To get exactly the <td> elements that have no other <td> elements as descendants (your "deepest" cells), you can use this XPath 1.0 expression, which works perfectly with lxml:

//td[not(descendant::td)]

How this works:

  • //td selects every <td> element in the entire document.
  • The predicate [not(descendant::td)] filters that list to only keep <td> elements where there are no <td> elements anywhere inside them (as children, grandchildren, etc.). This automatically excludes the outer <td> that wraps your nested table, since it contains a descendant <td>.

Example with lxml:

Here's how you'd use this in code to extract the desired cells from your sample HTML:

from lxml import etree

html_content = """
<table>
  <tr>
    <td>cell 1</td>
  </tr>
  <tr>
    <td><b>cell 2</b></td>
  </tr>
  <tr>
    <td>
      <table>
        <tr>
          <td><span>cell 3</span></td>
        </tr>
      </table>
    </td>
  </tr>
</table>
"""

tree = etree.HTML(html_content)
target_tds = tree.xpath('//td[not(descendant::td)]')

# Print the results
for td in target_tds:
    print(etree.tostring(td, encoding='unicode').strip())

Output:

<td>cell 1</td>
<td><b>cell 2</b></td>
<td><span>cell 3</span></td>

This gives you exactly the three <td> elements you wanted, excluding the outer one that contains the nested table.

内容的提问来源于stack exchange,提问作者sourcream

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 19:37:48