XPath查询求助:匹配非a标签包裹的首个wordToMatch实例
Got it, let's work through this XPath problem together. Your goal is to target the first occurrence of 'wordToMatch' that isn't wrapped in an <a> tag, and your current expression isn't delivering the right results. Let's break down what's wrong, then build the correct query.
What's Wrong With Your Current Expression?
Your attempt //text()[1][contains(.,'wordToMatch') and not(self::a)] has two key issues:
self::achecks if the text node itself is an<a>element—which it never is. Text nodes are children of elements, so you need to check their parent instead.//text()[1]selects the first text node under every ancestor node, not the first text node globally that meets your criteria.
The Correct XPath Expression
Here's the query that will do exactly what you need:
(//text()[contains(., 'wordToMatch') and not(parent::a)])[1]
Let's break this down:
//text(): Grabs all text nodes in the document.contains(., 'wordToMatch'): Filters those nodes to only ones that include your target text.not(parent::a): Excludes any text nodes whose direct parent is an<a>tag (i.e., text wrapped in<a>).(...)[1]: Wraps the filtered node set and selects the first node in that set—this ensures you get the very first occurrence that isn't wrapped in<a>.
Testing Against Your Examples
Let's verify this works with your scenarios:
- Example 1: The
<a>-wrappedwordToMatchis excluded. The first valid match is the one in the main<p>text. - Example 2: The
<b>-wrappedwordToMatchis included (since its parent is<b>, not<a>) and is the first valid occurrence. - Example 3: The second
wordToMatch(not in<a>) is the first valid hit, so it's selected, and later occurrences are ignored.
Java Implementation
Since you're working in Java, here's how to use this XPath with both standard XML parsing (for well-formed HTML/XML) and Jsoup (better for real-world HTML):
Option 1: Standard JAXP XML Parser
import javax.xml.parsers.DocumentBuilder; import javax.xml.parsers.DocumentBuilderFactory; import javax.xml.xpath.XPath; import javax.xml.xpath.XPathFactory; import org.w3c.dom.Document; import org.w3c.dom.Node; import java.io.ByteArrayInputStream; public class XPathMatcher { public static void main(String[] args) throws Exception { String html = "<p>Sample 1 <a href=\"shouldNotMatchWrappedInA\">wordToMatch</a> some random text to not be matched followed by wordToMatch, this should work.</p>"; // Initialize parser and document DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance(); DocumentBuilder builder = factory.newDocumentBuilder(); Document doc = builder.parse(new ByteArrayInputStream(html.getBytes())); // Execute XPath query XPath xpath = XPathFactory.newInstance().newXPath(); String query = "(//text()[contains(., 'wordToMatch') and not(parent::a)])[1]"; Node match = (Node) xpath.evaluate(query, doc, javax.xml.xpath.XPathConstants.NODE); if (match != null) { System.out.println("Matched text: " + match.getNodeValue().trim()); } else { System.out.println("No valid match found."); } } }
Option 2: Jsoup (Recommended for HTML)
Real-world HTML is often malformed, so Jsoup is a better choice. Here's how to use it with XPath:
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.select.XPathEvaluator; import org.jsoup.nodes.Node; public class JsoupXPathMatcher { public static void main(String[] args) { String html = "<p>Sample 2 <a href=\"shouldNotMatchWrappedInA\">wordToMatch</a> some random text to not be matched followed by <b>wordToMatch</b> this should work.</p>"; Document doc = Jsoup.parse(html); XPathEvaluator xpath = new XPathEvaluator(doc); Node match = xpath.evaluateFirst("(//text()[contains(., 'wordToMatch') and not(parent::a)])[1]"); if (match != null) { System.out.println("Matched text: " + match.toString().trim()); } else { System.out.println("No valid match found."); } } }
That should solve your problem perfectly! Let me know if you need any adjustments.
内容的提问来源于stack exchange,提问作者CoDemystified JavaFx

