XPath技术咨询:如何筛选首个指定class的div下不含关键词的文本节点
Let's break down what's going wrong with your current expression and build the correct one for your use case.
What's Wrong With Your Current XPath?
Your expression (//div[@class='blabla'])[1][not(contains(text(), 'bananas'))] has two key issues:
text()on the parentdivonly targets text nodes directly inside the div itself—not the text inside its childptags. Since your "bananas" text lives in a nestedp, this condition never actually checks for the keyword you care about.- You're filtering the entire parent
divinstead of the individual child nodes you want to exclude. Even if the condition worked, it would discard the entire firstblabladiv if any child had "bananas", which isn't what you want (you just want to skip the problematic child node, not the whole container).
The Correct XPath for Your Example
To get only the text from p tags inside the first blabla div, excluding any p that contains "bananas", use this expression:
(//div[@class='blabla'])[1]/p[not(contains(., 'bananas'))]/text()
Let's Break It Down:
(//div[@class='blabla'])[1]: Pinpoints the first div with classblabla(avoids matching the nested innerblabladiv that's a child of the first one)./p: Targets only direct childptags of that first div (if you wanted all descendantptags, use//pinstead—but your example doesn't include the inner div's text, so direct children align with your desired result).[not(contains(., 'bananas'))]: Filters out anyptag where the node's content (.refers to the currentpnode's full text) includes the word "bananas"./text(): Extracts the raw text content from the matchingptags.
If You Need All Descendant Text (Including Nested Nodes, Excluding Keywords)
If you ever want to capture all text inside the first blabla div (not just direct child ps) but still exclude any node that contains "bananas", use this:
(//div[@class='blabla'])[1]//*[not(contains(., 'bananas'))]/text()
This will grab text from any descendant element, as long as that element doesn't contain the excluded keyword.
Testing With Your Sample HTML
For your given markup:
<div class="blabla"> <p>I like bananas</p> <p>I also like apples</p> <div class="blabla"> <p>some text</p> </div> </div>
The first XPath will return exactly:
I also like apples
Which matches your desired result.
内容的提问来源于stack exchange,提问作者Adrien

