如何在Java中解析包含混合节点与文本的XML?
Got it, let's break down how to fix this. The problem with your initial XPath call is that it only grabs the text nodes inside the <sentence> tag, completely skipping over the <a> elements. To get the full mix of text and <a> elements, here are a few reliable approaches you can use:
1. Manually iterate over all child nodes
This gives you full control over how you combine text and <a> elements. First, grab the <sentence> element itself, then loop through every child node (text nodes and <a> elements) and build your content step by step:
// Get the <sentence> element directly Element sentenceElement = (Element) xPath.evaluate("sentence", transUnitElement, XPathConstants.NODE); StringBuilder fullContent = new StringBuilder(); // Loop through all child nodes (text + <a> elements) NodeList childNodes = sentenceElement.getChildNodes(); for (int i = 0; i < childNodes.getLength(); i++) { Node node = childNodes.item(i); if (node.getNodeType() == Node.TEXT_NODE) { // Append raw text content (trim if you want to clean up whitespace) fullContent.append(node.getTextContent().trim()); } else if (node.getNodeType() == Node.ELEMENT_NODE && "a".equals(node.getNodeName())) { // Handle the <a> element - include its tags, text, and any attributes you need Element aElement = (Element) node; fullContent.append("<a"); // Add attributes like href if present if (aElement.hasAttribute("href")) { fullContent.append(" href='").append(aElement.getAttribute("href")).append("'"); } fullContent.append(">").append(aElement.getTextContent()).append("</a>"); } } String completeSentence = fullContent.toString();
2. Serialize the entire <sentence> element
If you want to preserve the exact XML structure of the <sentence> content (including all <a> tags and their attributes), use a Transformer to convert the element to a string directly:
Element sentenceElement = (Element) xPath.evaluate("sentence", transUnitElement, XPathConstants.NODE); TransformerFactory transformerFactory = TransformerFactory.newInstance(); Transformer transformer = transformerFactory.newTransformer(); // Optional: Tweak output settings to remove extra whitespace/XML declaration transformer.setOutputProperty(OutputKeys.INDENT, "no"); transformer.setOutputProperty(OutputKeys.OMIT_XML_DECLARATION, "yes"); StringWriter writer = new StringWriter(); transformer.transform(new DOMSource(sentenceElement), new StreamResult(writer)); String completeContent = writer.toString();
This method is great if you don't need to modify the structure of the <a> elements and just want the full, intact content.
3. Use XPath to fetch all child nodes and process them
Combine XPath with element serialization to handle all child nodes in one go. This is useful if you want to target specific node types more precisely:
NodeList allChildNodes = (NodeList) xPath.evaluate("sentence/node()", transUnitElement, XPathConstants.NODESET); StringBuilder fullContent = new StringBuilder(); TransformerFactory tf = TransformerFactory.newInstance(); for (int i = 0; i < allChildNodes.getLength(); i++) { Node node = allChildNodes.item(i); if (node.getNodeType() == Node.TEXT_NODE) { fullContent.append(node.getTextContent().trim()); } else if (node.getNodeType() == Node.ELEMENT_NODE) { // Serialize any element (like <a>) to its XML string form Transformer t = tf.newTransformer(); t.setOutputProperty(OutputKeys.OMIT_XML_DECLARATION, "yes"); StringWriter sw = new StringWriter(); t.transform(new DOMSource(node), new StreamResult(sw)); fullContent.append(sw.toString()); } } String completeContent = fullContent.toString();
Pick the approach that fits your needs: manual iteration gives you maximum flexibility for custom formatting, while serialization is the quickest way to get the full XML structure.
内容的提问来源于stack exchange,提问作者Pratik

