Java提取XML标签与属性间文本的技术问题求助
Since GEXF is an XML-based format, you have two main approaches to extract content between <nodes> tags: a quick regex solution (for simple cases) or a robust XML parser (recommended for production use). Let’s break both down with complete code examples.
Regex Approach (Quick but Limited)
This works if your XML is well-formed, has no nested <nodes> tags, and no edge cases like comments or CDATA containing </nodes>.
Here’s the full implementation:
import java.io.BufferedReader; import java.io.FileReader; import java.io.IOException; import java.util.regex.Matcher; import java.util.regex.Pattern; public class Main { private static String filePath = "src/babel.gexf"; public static void main(String[] args) { String fileContent = readFileToString(filePath); if (fileContent != null) { String nodesContent = extractNodesWithRegex(fileContent); System.out.println("Content inside <nodes>:\n" + nodesContent); } } // Read the entire GEXF file into a string private static String readFileToString(String filePath) { StringBuilder content = new StringBuilder(); try (BufferedReader br = new BufferedReader(new FileReader(filePath))) { String line; while ((line = br.readLine()) != null) { content.append(line).append("\n"); } } catch (IOException e) { e.printStackTrace(); return null; } return content.toString(); } // Extract content between <nodes> and </nodes> using regex private static String extractNodesWithRegex(String xmlContent) { // Use DOTALL to make . match newlines, non-greedy *? to stop at first </nodes> Pattern pattern = Pattern.compile("<nodes>(.*?)</nodes>", Pattern.DOTALL); Matcher matcher = pattern.matcher(xmlContent); if (matcher.find()) { return matcher.group(1).trim(); // Trim extra whitespace/newlines } return "<nodes> tag not found"; } }
Key Regex Details:
Pattern.DOTALL: Ensures the.character matches newline characters (critical since XML content spans multiple lines)..*?: Non-greedy quantifier that stops at the first closing</nodes>tag (avoids capturing everything up to the last closing tag if there were duplicates).
XML Parser Approach (Recommended)
Regex can break easily with unexpected XML formatting (like nested tags, comments, or CDATA). Using a dedicated XML parser is far more reliable. Here’s how to do it with Java’s built-in DOM parser:
import org.w3c.dom.Document; import org.w3c.dom.Node; import org.w3c.dom.NodeList; import javax.xml.parsers.DocumentBuilder; import javax.xml.parsers.DocumentBuilderFactory; import java.io.File; public class Main { private static String filePath = "src/babel.gexf"; public static void main(String[] args) { try { DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance(); DocumentBuilder builder = factory.newDocumentBuilder(); Document doc = builder.parse(new File(filePath)); // Get the first <nodes> element Node nodesElement = doc.getElementsByTagName("nodes").item(0); if (nodesElement != null) { StringBuilder nodesContent = new StringBuilder(); NodeList childNodes = nodesElement.getChildNodes(); // Iterate through all child content (tags and text) for (int i = 0; i < childNodes.getLength(); i++) { Node child = childNodes.item(i); // Append raw text content (use a Transformer if you need full XML tags) nodesContent.append(child.getTextContent()).append("\n"); } System.out.println("Content inside <nodes>:\n" + nodesContent.toString().trim()); } else { System.out.println("<nodes> tag not found in the file"); } } catch (Exception e) { e.printStackTrace(); } } }
Why This Is Better:
- Handles all XML edge cases (nested tags, comments, CDATA, malformed XML warnings).
- Lets you easily access individual
<node>elements inside<nodes>later if you need to process them further. - Follows XML best practices and is maintainable for larger projects.
Final Notes:
- Use regex only for throwaway scripts or strictly controlled XML scenarios.
- For production code, always use an XML parser (DOM, SAX, or a library like JAXB) to avoid brittle code.
内容的提问来源于stack exchange,提问作者rawsly
相关产品推荐
相关产品推荐

