C#使用XmlDocument解析XML时如何获取可选annotation标签
C# 解析XML可选annotation节点实现方案
核心逻辑:<annotation>是<sample>节点下的可选直接子元素,通过XPath定位子节点后判断存在性,分别对应赋值规则即可。
完整实现代码
XmlDocument xdoc = new XmlDocument(); xdoc.Load(fileName); XmlNodeList nodes = xdoc.SelectNodes("project/samples/sample"); foreach (XmlNode xn in nodes) { ParseList tempList = new ParseList(); tempList.intentName = xn.Attributes["intentref"]?.Value ?? "Unassigned"; tempList.sampleSentence = xn.FirstChild.Value; tempList.countSamples = xn.Attributes["count"].Value; // 查找当前sample节点下的annotation直接子节点 XmlNode annotationNode = xn.SelectSingleNode("annotation"); if (annotationNode == null) { tempList.annotatedText = "NONE"; } else { // 读取annotation完整标签内容,匹配示例输出格式 string rawAnno = annotationNode.OuterXml.Replace("\"o_number\"", "o_number"); tempList.annotatedText = $"\"{rawAnno}\""; } }
实现说明
SelectSingleNode("annotation")只会匹配当前sample节点下的直接子级annotation标签,不会递归遍历深层节点,完全匹配当前XML结构- 节点的
OuterXml属性会直接返回完整的标签字符串,包含标签名、属性、内部文本,不需要手动拼接结构 - 如果不需要匹配示例中去掉属性值双引号的效果,直接删除
.Replace("\"o_number\"", "o_number")即可,输出会和原始XML的annotation标签格式完全一致 - 如果后续出现单sample多annotation的场景,替换
SelectSingleNode为SelectNodes("annotation")遍历拼接所有annotation内容即可
注:当前
sampleSentence的取值写法在你给出的XML结构下可以正常运行,如果后续出现annotation标签插在句子中间的标注格式,建议改为提取sample下所有文本节点的内容拼接,避免取到不完整的句子。
内容的提问来源于stack exchange,提问作者aarong_
相关产品推荐
相关产品推荐

