You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java中如何简洁高效地将XML转换为不含属性的JSON?

Efficiently Convert XML to JSON Without Attributes

Absolutely agree that post-processing a full JSONObject to strip attributes is clunky and inefficient—you’re way better off handling this during the XML parsing step instead of cleaning up after the fact. Here are some practical, efficient solutions depending on your tech stack:

Python Solutions

Using xmltodict (Simplest Approach)

The xmltodict library makes this trivial with a custom postprocessor that filters out attribute keys (which it prefixes with @ by default). This avoids loading unnecessary attribute data into memory entirely.

First install the library:

pip install xmltodict

Then use it with a postprocessor to drop attributes:

import xmltodict

def filter_attributes(path, key, value):
    # Skip any keys that represent XML attributes (prefixed with @)
    if key.startswith('@'):
        return None
    return key, value

xml_input = """<root id="120"> <child1 id="21">val1</child1> <child2 id="22">val2</child2> </root>"""
# Parse with our attribute-filtering postprocessor
result = xmltodict.parse(
    xml_input,
    postprocessor=filter_attributes,
    dict_constructor=dict
)

print(result)
# Output: {'root': {'child1': 'val1', 'child2': 'val2'}}

SAX Parsing (Best for Large Files)

If you’re working with huge XML files and want minimal memory usage, use Python’s built-in SAX parser. It streams the XML instead of loading the entire tree into memory, and you can simply ignore attributes as you parse.

import xml.sax
import json

class AttributeStrippingHandler(xml.sax.ContentHandler):
    def __init__(self):
        self.data = {}
        self.current_element_stack = []
        self.current_text = ""

    def startElement(self, name, attrs):
        # Ignore attrs completely, just track the element path
        self.current_element_stack.append(name)
        self.current_text = ""

    def characters(self, content):
        # Capture and clean up text content
        self.current_text += content.strip()

    def endElement(self, name):
        # Traverse to the parent element in our data structure
        current_node = self.data
        for elem in self.current_element_stack[:-1]:
            if elem not in current_node:
                current_node[elem] = {}
            current_node = current_node[elem]
        
        # Set the text value if there's content (handles leaf nodes)
        if self.current_text:
            current_node[name] = self.current_text
        
        self.current_element_stack.pop()

# Usage
xml_input = """<root id="120"> <child1 id="21">val1</child1> <child2 id="22">val2</child2> </root>"""
handler = AttributeStrippingHandler()
xml.sax.parseString(xml_input, handler)

print(json.dumps(handler.data, indent=2))
# Output:
# {
#   "root": {
#     "child1": "val1",
#     "child2": "val2"
#   }
# }

Java Solution (Using Jackson)

Jackson’s XmlMapper has a built-in feature to ignore all XML attributes during parsing, which is far more efficient than post-processing.

First add the Jackson XML dependency to your project (Maven example):

<dependency>
    <groupId>com.fasterxml.jackson.dataformat</groupId>
    <artifactId>jackson-dataformat-xml</artifactId>
    <version>2.15.2</version>
</dependency>

Then configure the mapper to ignore attributes:

import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.dataformat.xml.XmlMapper;
import com.fasterxml.jackson.dataformat.xml.deser.FromXmlParser;

public class XmlToJsonWithoutAttributes {
    public static void main(String[] args) throws Exception {
        String xmlInput = "<root id=\"120\"> <child1 id=\"21\">val1</child1> <child2 id=\"22\">val2</child2> </root>";
        
        XmlMapper xmlMapper = new XmlMapper();
        // Enable the feature to ignore all XML attributes
        xmlMapper.configure(FromXmlParser.Feature.IGNORE_ATTRIBUTES, true);
        
        JsonNode result = xmlMapper.readTree(xmlInput);
        System.out.println(result.toPrettyString());
        // Output:
        // {
        //   "root" : {
        //     "child1" : "val1",
        //     "child2" : "val2"
        //   }
        // }
    }
}

Core Principle

The key to efficiency here is ignoring attributes at the parsing stage, not after converting to a JSON structure. This avoids wasting memory on data you don’t need and skips the extra step of traversing the JSON to delete attributes later.

内容的提问来源于stack exchange,提问作者user3243499

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:32:16