You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解析含自闭合Node标签的XML并提取纯文本(保留索引)

提取XML中的纯文本(跳过自闭合Node标签)

嘿,这个问题我太懂了!要保留Node标签的索引信息但只提取纯文本,其实不用复杂操作,根据你使用的编程语言,这几种方法都能轻松搞定:

方案1:Python(用标准XML解析库)

Python自带的xml.etree.ElementTree就能完美处理,它会自动跳过所有子元素节点,只提取文本内容:

import xml.etree.ElementTree as ET

# 你的XML片段
xml_content = '''<TextWithNodes> <Node id="0"/>A TEENAGER <Node id="11"/>yesterday<Node id="20"/> accused his parents of cruelty by feeding him a daily diet of chips which sent his weight ballooning...</TextWithNodes>'''

# 解析XML
root = ET.fromstring(xml_content)
# 遍历所有文本节点并拼接
pure_text = ''.join(root.itertext()).strip()
print(pure_text)

输出结果就是:A TEENAGER yesterday accused his parents of cruelty by feeding him a daily diet of chips which sent his weight ballooning...

方案2:Java(DOM解析)

Java的DOM API里的getTextContent()方法直接就能获取元素下所有纯文本,完全忽略中间的标签:

import org.w3c.dom.Document;
import org.w3c.dom.Element;
import javax.xml.parsers.DocumentBuilder;
import javax.xml.parsers.DocumentBuilderFactory;
import java.io.ByteArrayInputStream;

public class ExtractXmlText {
    public static void main(String[] args) throws Exception {
        String xml = "<TextWithNodes> <Node id=\"0\"/>A TEENAGER <Node id=\"11\"/>yesterday<Node id=\"20\"/> accused his parents of cruelty by feeding him a daily diet of chips which sent his weight ballooning...</TextWithNodes>";
        
        DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
        DocumentBuilder builder = factory.newDocumentBuilder();
        Document doc = builder.parse(new ByteArrayInputStream(xml.getBytes()));
        
        // 获取TextWithNodes元素
        Element textContainer = (Element) doc.getElementsByTagName("TextWithNodes").item(0);
        // 提取纯文本
        String pureText = textContainer.getTextContent().trim();
        System.out.println(pureText);
    }
}

应急方案:正则表达式(不推荐复杂XML场景)

如果只是临时处理结构非常固定的XML片段,正则可以应急,但不推荐用于复杂XML(比如Node标签有嵌套、属性含特殊字符时会失效):

import re

xml_content = '''<TextWithNodes> <Node id="0"/>A TEENAGER <Node id="11"/>yesterday<Node id="20"/> accused his parents of cruelty by feeding him a daily diet of chips which sent his weight ballooning...</TextWithNodes>'''

# 匹配所有自闭合的Node标签并替换为空
pure_text = re.sub(r'<Node[^/>]+/>', '', xml_content)
# 再去掉TextWithNodes的首尾标签并清理空格
pure_text = pure_text.replace('<TextWithNodes>', '').replace('</TextWithNodes>', '').strip()
print(pure_text)

内容的提问来源于stack exchange,提问作者Elliott

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:09:42