如何解析含自闭合Node标签的XML并提取纯文本(保留索引)
提取XML中的纯文本(跳过自闭合Node标签)
嘿,这个问题我太懂了!要保留Node标签的索引信息但只提取纯文本,其实不用复杂操作,根据你使用的编程语言,这几种方法都能轻松搞定:
方案1:Python(用标准XML解析库)
Python自带的xml.etree.ElementTree就能完美处理,它会自动跳过所有子元素节点,只提取文本内容:
import xml.etree.ElementTree as ET # 你的XML片段 xml_content = '''<TextWithNodes> <Node id="0"/>A TEENAGER <Node id="11"/>yesterday<Node id="20"/> accused his parents of cruelty by feeding him a daily diet of chips which sent his weight ballooning...</TextWithNodes>''' # 解析XML root = ET.fromstring(xml_content) # 遍历所有文本节点并拼接 pure_text = ''.join(root.itertext()).strip() print(pure_text)
输出结果就是:A TEENAGER yesterday accused his parents of cruelty by feeding him a daily diet of chips which sent his weight ballooning...
方案2:Java(DOM解析)
Java的DOM API里的getTextContent()方法直接就能获取元素下所有纯文本,完全忽略中间的
import org.w3c.dom.Document; import org.w3c.dom.Element; import javax.xml.parsers.DocumentBuilder; import javax.xml.parsers.DocumentBuilderFactory; import java.io.ByteArrayInputStream; public class ExtractXmlText { public static void main(String[] args) throws Exception { String xml = "<TextWithNodes> <Node id=\"0\"/>A TEENAGER <Node id=\"11\"/>yesterday<Node id=\"20\"/> accused his parents of cruelty by feeding him a daily diet of chips which sent his weight ballooning...</TextWithNodes>"; DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance(); DocumentBuilder builder = factory.newDocumentBuilder(); Document doc = builder.parse(new ByteArrayInputStream(xml.getBytes())); // 获取TextWithNodes元素 Element textContainer = (Element) doc.getElementsByTagName("TextWithNodes").item(0); // 提取纯文本 String pureText = textContainer.getTextContent().trim(); System.out.println(pureText); } }
应急方案:正则表达式(不推荐复杂XML场景)
如果只是临时处理结构非常固定的XML片段,正则可以应急,但不推荐用于复杂XML(比如Node标签有嵌套、属性含特殊字符时会失效):
import re xml_content = '''<TextWithNodes> <Node id="0"/>A TEENAGER <Node id="11"/>yesterday<Node id="20"/> accused his parents of cruelty by feeding him a daily diet of chips which sent his weight ballooning...</TextWithNodes>''' # 匹配所有自闭合的Node标签并替换为空 pure_text = re.sub(r'<Node[^/>]+/>', '', xml_content) # 再去掉TextWithNodes的首尾标签并清理空格 pure_text = pure_text.replace('<TextWithNodes>', '').replace('</TextWithNodes>', '').strip() print(pure_text)
内容的提问来源于stack exchange,提问作者Elliott
相关产品推荐
相关产品推荐

