You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java提取XML标签与属性间文本的技术问题求助

Extracting Text Between <nodes> Tags in a GEXF File with Java

Since GEXF is an XML-based format, you have two main approaches to extract content between <nodes> tags: a quick regex solution (for simple cases) or a robust XML parser (recommended for production use). Let’s break both down with complete code examples.

Regex Approach (Quick but Limited)

This works if your XML is well-formed, has no nested <nodes> tags, and no edge cases like comments or CDATA containing </nodes>.

Here’s the full implementation:

import java.io.BufferedReader;
import java.io.FileReader;
import java.io.IOException;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class Main {
    private static String filePath = "src/babel.gexf";

    public static void main(String[] args) {
        String fileContent = readFileToString(filePath);
        if (fileContent != null) {
            String nodesContent = extractNodesWithRegex(fileContent);
            System.out.println("Content inside <nodes>:\n" + nodesContent);
        }
    }

    // Read the entire GEXF file into a string
    private static String readFileToString(String filePath) {
        StringBuilder content = new StringBuilder();
        try (BufferedReader br = new BufferedReader(new FileReader(filePath))) {
            String line;
            while ((line = br.readLine()) != null) {
                content.append(line).append("\n");
            }
        } catch (IOException e) {
            e.printStackTrace();
            return null;
        }
        return content.toString();
    }

    // Extract content between <nodes> and </nodes> using regex
    private static String extractNodesWithRegex(String xmlContent) {
        // Use DOTALL to make . match newlines, non-greedy *? to stop at first </nodes>
        Pattern pattern = Pattern.compile("<nodes>(.*?)</nodes>", Pattern.DOTALL);
        Matcher matcher = pattern.matcher(xmlContent);
        
        if (matcher.find()) {
            return matcher.group(1).trim(); // Trim extra whitespace/newlines
        }
        return "<nodes> tag not found";
    }
}

Key Regex Details:

  • Pattern.DOTALL: Ensures the . character matches newline characters (critical since XML content spans multiple lines).
  • .*?: Non-greedy quantifier that stops at the first closing </nodes> tag (avoids capturing everything up to the last closing tag if there were duplicates).

Regex can break easily with unexpected XML formatting (like nested tags, comments, or CDATA). Using a dedicated XML parser is far more reliable. Here’s how to do it with Java’s built-in DOM parser:

import org.w3c.dom.Document;
import org.w3c.dom.Node;
import org.w3c.dom.NodeList;
import javax.xml.parsers.DocumentBuilder;
import javax.xml.parsers.DocumentBuilderFactory;
import java.io.File;

public class Main {
    private static String filePath = "src/babel.gexf";

    public static void main(String[] args) {
        try {
            DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
            DocumentBuilder builder = factory.newDocumentBuilder();
            Document doc = builder.parse(new File(filePath));
            
            // Get the first <nodes> element
            Node nodesElement = doc.getElementsByTagName("nodes").item(0);
            if (nodesElement != null) {
                StringBuilder nodesContent = new StringBuilder();
                NodeList childNodes = nodesElement.getChildNodes();
                
                // Iterate through all child content (tags and text)
                for (int i = 0; i < childNodes.getLength(); i++) {
                    Node child = childNodes.item(i);
                    // Append raw text content (use a Transformer if you need full XML tags)
                    nodesContent.append(child.getTextContent()).append("\n");
                }
                
                System.out.println("Content inside <nodes>:\n" + nodesContent.toString().trim());
            } else {
                System.out.println("<nodes> tag not found in the file");
            }
        } catch (Exception e) {
            e.printStackTrace();
        }
    }
}

Why This Is Better:

  • Handles all XML edge cases (nested tags, comments, CDATA, malformed XML warnings).
  • Lets you easily access individual <node> elements inside <nodes> later if you need to process them further.
  • Follows XML best practices and is maintainable for larger projects.

Final Notes:

  • Use regex only for throwaway scripts or strictly controlled XML scenarios.
  • For production code, always use an XML parser (DOM, SAX, or a library like JAXB) to avoid brittle code.

内容的提问来源于stack exchange,提问作者rawsly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:33:40