C# XmlReader无标签间空格时忽略标签问题排查
问题:XmlReader读取单行XML时跳过标签的异常
我用同一程序内的XmlWriter输出了单行XML,内容如下:
<route><name>CSX Atlanta Division Abbeville Sub</name><description>The Abbeville Subdivision runs from Tucker, GA to Abbeville, SC. It is a CTC-line, single track with passing sidings. There are four road trains, coal, grain, ethanol, and local traffic.</description><history>The Abbeville Subdivision was originally part of the Seaboard Air Line Railroad. With the creation of CSX in the early 1980s, the Abbeville Sub has remained a part of CSX since.</history><radio></radio></route>
但使用XmlReader读取时出现异常:读取完name标签后直接跳过description标签,读取history标签。读取代码及常量定义如下:
public class Constants { public class Tags { public const string NAME = "name"; public const string DESC = "description"; public const string NOTES = "notes"; public const string TYPE = "type"; public class RouteTags { public const string ROUTE = "route"; public const string HISTORY = "history"; public const string HISTORYLINK = "historylink"; public const string INDUSTRY = "industry"; public const string LOCATIONINFO = "locationinfo"; } } public void ReadXML(XmlReader textReader) { while (textReader.Read()) { if(textReader.Name == Constants.Tags.NAME && textReader.IsStartElement()) { Name = textReader.ReadElementContentAsString(); } else if (textReader.Name == Constants.Tags.DESC && textReader.IsStartElement()) { Description = textReader.ReadElementContentAsString(); } else if (textReader.Name == Constants.Tags.RouteTags.HISTORY && textReader.IsStartElement()) { History = textReader.ReadElementContentAsString(); } else if(textReader.Name == TAG && !textReader.IsStartElement()) { return; } } }
疑问:为何XmlReader不识别name的结束标签并跳转至description的开始标签?目前仅能通过手动添加标签间空格或格式化XML正常读取,但程序要求无需人工干预即可读取自身输出的XML,该如何解决?
原因及解决办法
问题根源
ReadElementContentAsString()方法会自动将XmlReader的指针移动到当前元素结束标签的之后位置。当你读取完name元素后,指针已经定位到<description>的开始标签上,但下一次循环调用Read()会直接跳过这个开始标签,移动到description的文本内容节点——此时textReader.Name虽然还是description,但IsStartElement()会返回false,导致无法进入对应的处理分支,最终直接跳过该元素,直到遇到<history>的开始标签。
解决方法一:调整XmlReader读取逻辑
改用MoveToContent()替代循环开头的Read(),它会自动跳过空白、注释等无效节点,直接定位到有效内容节点,避免指针重复移动导致的元素跳过:
public void ReadXML(XmlReader textReader) { while (textReader.MoveToContent() != XmlNodeType.EndElement) { if (textReader.IsStartElement(Constants.Tags.NAME)) { Name = textReader.ReadElementContentAsString(); } else if (textReader.IsStartElement(Constants.Tags.DESC)) { Description = textReader.ReadElementContentAsString(); } else if (textReader.IsStartElement(Constants.Tags.RouteTags.HISTORY)) { History = textReader.ReadElementContentAsString(); } else { // 跳过无需处理的元素(比如示例中的<radio>) textReader.Skip(); } } }
解决方法二:改用LINQ to XML(更简洁可靠)
如果业务允许,直接使用LINQ to XML可以彻底避免手动处理XmlReader指针的问题,代码可读性更高,且对单行/格式化XML都能正常解析:
public void ReadXML(XmlReader textReader) { XElement route = XElement.Load(textReader); Name = route.Element(Constants.Tags.NAME)?.Value; Description = route.Element(Constants.Tags.DESC)?.Value; History = route.Element(Constants.Tags.RouteTags.HISTORY)?.Value; }
内容的提问来源于stack exchange,提问作者MattCW
相关产品推荐
相关产品推荐

