Java中如何简洁高效地将XML转换为不含属性的JSON?
Absolutely agree that post-processing a full JSONObject to strip attributes is clunky and inefficient—you’re way better off handling this during the XML parsing step instead of cleaning up after the fact. Here are some practical, efficient solutions depending on your tech stack:
Python Solutions
Using xmltodict (Simplest Approach)
The xmltodict library makes this trivial with a custom postprocessor that filters out attribute keys (which it prefixes with @ by default). This avoids loading unnecessary attribute data into memory entirely.
First install the library:
pip install xmltodict
Then use it with a postprocessor to drop attributes:
import xmltodict def filter_attributes(path, key, value): # Skip any keys that represent XML attributes (prefixed with @) if key.startswith('@'): return None return key, value xml_input = """<root id="120"> <child1 id="21">val1</child1> <child2 id="22">val2</child2> </root>""" # Parse with our attribute-filtering postprocessor result = xmltodict.parse( xml_input, postprocessor=filter_attributes, dict_constructor=dict ) print(result) # Output: {'root': {'child1': 'val1', 'child2': 'val2'}}
SAX Parsing (Best for Large Files)
If you’re working with huge XML files and want minimal memory usage, use Python’s built-in SAX parser. It streams the XML instead of loading the entire tree into memory, and you can simply ignore attributes as you parse.
import xml.sax import json class AttributeStrippingHandler(xml.sax.ContentHandler): def __init__(self): self.data = {} self.current_element_stack = [] self.current_text = "" def startElement(self, name, attrs): # Ignore attrs completely, just track the element path self.current_element_stack.append(name) self.current_text = "" def characters(self, content): # Capture and clean up text content self.current_text += content.strip() def endElement(self, name): # Traverse to the parent element in our data structure current_node = self.data for elem in self.current_element_stack[:-1]: if elem not in current_node: current_node[elem] = {} current_node = current_node[elem] # Set the text value if there's content (handles leaf nodes) if self.current_text: current_node[name] = self.current_text self.current_element_stack.pop() # Usage xml_input = """<root id="120"> <child1 id="21">val1</child1> <child2 id="22">val2</child2> </root>""" handler = AttributeStrippingHandler() xml.sax.parseString(xml_input, handler) print(json.dumps(handler.data, indent=2)) # Output: # { # "root": { # "child1": "val1", # "child2": "val2" # } # }
Java Solution (Using Jackson)
Jackson’s XmlMapper has a built-in feature to ignore all XML attributes during parsing, which is far more efficient than post-processing.
First add the Jackson XML dependency to your project (Maven example):
<dependency> <groupId>com.fasterxml.jackson.dataformat</groupId> <artifactId>jackson-dataformat-xml</artifactId> <version>2.15.2</version> </dependency>
Then configure the mapper to ignore attributes:
import com.fasterxml.jackson.databind.JsonNode; import com.fasterxml.jackson.dataformat.xml.XmlMapper; import com.fasterxml.jackson.dataformat.xml.deser.FromXmlParser; public class XmlToJsonWithoutAttributes { public static void main(String[] args) throws Exception { String xmlInput = "<root id=\"120\"> <child1 id=\"21\">val1</child1> <child2 id=\"22\">val2</child2> </root>"; XmlMapper xmlMapper = new XmlMapper(); // Enable the feature to ignore all XML attributes xmlMapper.configure(FromXmlParser.Feature.IGNORE_ATTRIBUTES, true); JsonNode result = xmlMapper.readTree(xmlInput); System.out.println(result.toPrettyString()); // Output: // { // "root" : { // "child1" : "val1", // "child2" : "val2" // } // } } }
Core Principle
The key to efficiency here is ignoring attributes at the parsing stage, not after converting to a JSON structure. This avoids wasting memory on data you don’t need and skips the extra step of traversing the JSON to delete attributes later.
内容的提问来源于stack exchange,提问作者user3243499

