如何使用XPath 1.0获取最深层的<td>元素(适配lxml库)
Selecting Deepest Elements (No Descendant ) with XPath 1.0
To get exactly the <td> elements that have no other <td> elements as descendants (your "deepest" cells), you can use this XPath 1.0 expression, which works perfectly with lxml:
//td[not(descendant::td)]
How this works:
//tdselects every<td>element in the entire document.- The predicate
[not(descendant::td)]filters that list to only keep<td>elements where there are no<td>elements anywhere inside them (as children, grandchildren, etc.). This automatically excludes the outer<td>that wraps your nested table, since it contains a descendant<td>.
Example with lxml:
Here's how you'd use this in code to extract the desired cells from your sample HTML:
from lxml import etree html_content = """ <table> <tr> <td>cell 1</td> </tr> <tr> <td><b>cell 2</b></td> </tr> <tr> <td> <table> <tr> <td><span>cell 3</span></td> </tr> </table> </td> </tr> </table> """ tree = etree.HTML(html_content) target_tds = tree.xpath('//td[not(descendant::td)]') # Print the results for td in target_tds: print(etree.tostring(td, encoding='unicode').strip())
Output:
<td>cell 1</td> <td><b>cell 2</b></td> <td><span>cell 3</span></td>
This gives you exactly the three <td> elements you wanted, excluding the outer one that contains the nested table.
内容的提问来源于stack exchange,提问作者sourcream
相关产品推荐
相关产品推荐

