为type="float"的fig元素生成ID并封装文本到caption的技术方案
处理XML中float类型fig元素的ID生成与内容重构
需求说明
需要完成两项操作:
- 为所有
type="float"的<fig>元素生成格式为fg+数字序号的唯一ID - 将
<fig>元素内的文本节点转移到新创建的<caption>子元素中
现有C#代码无法正确筛选目标元素并完成上述操作,需修正代码以达成目标XML结构。
输入XML
<article> <fm> <journalTitle>The j Title</journalTitle> </fm> <body> <sec1> <p>The first para <fig type="float">This is First float image</fig></p> <p>THe other text <fig type="inline">This is First Inline image</fig></p> <sec2> <p>The first para <fig type="float">This is Second float image</fig></p> <p>THe other text <fig type="inline">This is First Inline image</fig></p> <sec3> <p>The first para <fig type="float">This is third float image</fig></p> <p>THe other text</p> </sec3> </sec2> </sec1> </body> </article>
现有C#代码
using System; using System.Collections.Generic; using System.Linq; using System.Xml.Linq; using System.Text.RegularExpressions; using System.Xml.XPath; namespace FigSeq1 { class Program { static void Main(string[] args) { XDocument XDoc1 = new XDocument(); XDoc1 = XDocument.Load(@"D:\XMLs\InImages.xml"); //Grouping only figs of type <fig type="float"> var FigGroup = XDoc1.Descendants() .Where(e=>e.Name.ToString()=="fig") .Attributes().Where(e1=>(e1.Name.ToString()=="type").ToString()=="float"); int cnt = 1; foreach(XElement ele1 in FigGroup) { ele1.Add(new XAttribute("ID", "fg" + cnt)); var EleAll = ele1.Nodes(); ele1.Nodes().Remove(); ele1.Add(new XElement("caption", EleAll)); cnt++; } XDoc1.Save(@"D:\XMLs\Result.xml"); } } }
期望输出XML
<article> <fm> <journalTitle>The j Title</journalTitle> </fm> <body> <sec1> <p>The first para <fig type="float" ID="fg1"><caption>This is First float image</caption></fig></p> <p>THe other text <fig type="inline">This is First Inline image</fig></p> <sec2> <p>The first para <fig type="float" ID="fg2"><caption>This is Second float image</caption></fig></p> <p>THe other text <fig type="inline">This is First Inline image</fig></p> <sec3> <p>The first para <fig type="float" ID="fg3"><caption>This is third float image</caption></fig></p> <p>THe other text</p> </sec3> </sec2> </sec1> </body> </article>
现有代码问题分析
- 筛选逻辑错误:原代码链式调用
.Attributes()后,得到的是符合条件的属性节点而非<fig>元素,后续循环操作对象错误,完全无法实现需求。 - 条件判断错误:
(e1.Name.ToString()=="type").ToString()=="float"逻辑混乱,实际是将布尔值转为字符串后和"float"比较,完全不符合预期。
修正后的C#代码
using System; using System.Linq; using System.Xml.Linq; namespace FigSeq1 { class Program { static void Main(string[] args) { // 加载目标XML文档 XDocument doc = XDocument.Load(@"D:\XMLs\InImages.xml"); // 精准筛选所有type属性为float的fig元素 var floatFigs = doc.Descendants("fig") .Where(fig => fig.Attribute("type")?.Value == "float"); int counter = 1; foreach (XElement fig in floatFigs) { // 生成并设置ID属性 fig.SetAttributeValue("ID", $"fg{counter}"); // 保存fig原有子节点,避免枚举时集合变更异常 var contentNodes = fig.Nodes().ToList(); // 清空fig原有内容 fig.RemoveNodes(); // 创建caption元素并移入原有内容 fig.Add(new XElement("caption", contentNodes)); counter++; } // 保存修改后的XML doc.Save(@"D:\XMLs\Result.xml"); } } }
修正说明
- 修复筛选逻辑:直接获取
<fig>元素,通过Attribute("type")?.Value == "float"判断目标元素,确保操作对象是<fig>节点。 - 优化空值处理:使用
?.空值运算符,避免因<fig>元素无type属性导致的空引用异常。 - 改进节点操作:用
ToList()固化节点集合,防止移除节点时引发枚举异常,确保内容正确转移到<caption>中。 - 精简代码:移除未使用的命名空间,简化变量命名,提升代码可读性。
内容的提问来源于stack exchange,提问作者Rudramuni TP
相关产品推荐
相关产品推荐

