如何使用CSS选择器选中除table元素及其内容外的所有元素?
选中div.page内元素并排除table及其内容的解决方案
你之前用的:not(table), :not(table) *选择器无效,是因为它会匹配到table元素的父容器(比如示例中的div.table)的子元素——table本身,因为父容器不是table,所以它的子元素table会被:not(table) *选中,导致table及其内容依然被包含。
下面提供两种可行的解决方案:
一、CSS选择器方案
使用更精准的选择器,排除table元素及其所有后代:
from requests_html import HTML html_string="""<html> <head> <title>Testing</title> </head> <body> <div class="page group"> <h2 class="level2">This is a heading</h2> <p>This is a paragraph.</p> <div class="table"> <table border=1> <tr class="row1"> <th>th 1</th> <td>td 1</td> </tr> <tr class="row2"> <th>th 2</th> <td>td 2</td> </tr> </table> </div> <h3 class="level3">a sub heading</h3> <p>This is also a paragraph.</p> <p>This is another paragraph.</p> <div>This is some text in a div element.</div> <a href="https://www.blah.com" target="_blank">Blah!</a> </div> </body>""" page = HTML(html=html_string) # 选中div.page下所有非table、且非table后代的元素 elements = page.find('div.page *:not(table):not(table *)') # 验证结果 for elem in elements: print(f"{elem.tag}: {elem.text.strip()}")
选择器逻辑:
div.page *:选中div.page下的所有子元素:not(table):排除table元素本身:not(table *):排除table的所有后代元素
如果需要连包含table的div.table容器一起排除,可改用选择器:div.page > *:not(.table)
二、XPath方案
XPath的逻辑更直观,适合复杂的排除场景:
from requests_html import HTML # 复用上面的html_string... page = HTML(html=html_string) # XPath表达式:选中div.page下所有不是table、且祖先中没有table的元素 elements = page.xpath('//div[contains(@class, "page")]//*[not(self::table) and not(ancestor::table)]') # 验证结果 for elem in elements: print(f"{elem.tag}: {elem.text.strip()}")
XPath逻辑:
//div[contains(@class, "page")]:定位到目标div.page容器//*:选中容器下的所有元素not(self::table):排除table元素本身not(ancestor::table):排除所有以table为祖先的元素(即table的所有后代)
内容的提问来源于stack exchange,提问作者Calvin Broadus
相关产品推荐
相关产品推荐

