如何跨两个项目符号匹配文本并创建内容控件(SDT)?
问题描述
假设文档包含以下内容:
- Hello
- World
需要编写正则表达式匹配"Hello"和"World"并创建内容控件(SDT),当前使用Aspose.Words 25.3.0版本,现有代码无法匹配跨两个列表项的文本,请问能否实现该需求?
现有代码如下:
string pattern = @"Hello\s*World"; var options = new FindReplaceOptions { ReplacingCallback = new ReplaceWithContentControlHandler(doc, "MyContentControl"), MatchCase = false, FindWholeWordsOnly = false, SmartParagraphBreakReplacement = true }; var patern = new Regex(pattern); doc.Range.Replace(new Regex(pattern, RegexOptions.IgnoreCase), "", options); class ReplaceWithContentControlHandler : IReplacingCallback { private readonly Document _doc; private readonly string _title; private readonly string _tag; public ReplaceWithContentControlHandler(Document doc, string title) { _doc = doc; _title = title; } public ReplaceAction Replacing(ReplacingArgs e) { // Create a new RichText content control StructuredDocumentTag sdt = new StructuredDocumentTag(_doc, SdtType.RichText, MarkupLevel.Inline) { LockContentControl = false, LockContents = false, Title = "Clause", IsShowingPlaceholderText = false }; sdt.RemoveAllChildren(); // Create a new Run node and add it to the content control Run run1 = new Run(_doc, e.Match.Value); sdt.AppendChild(run1); // Insert the content control into the document Node currentNode = e.MatchNode; // Check if the current node is a paragraph if (currentNode.NodeType == NodeType.Paragraph) { Paragraph currentParagraph = (Paragraph)currentNode; // If the match is at the beginning of the paragraph if (e.MatchOffset == 0) { // Create a new Run node and add it to the content control Run run = new Run(_doc, e.Match.Value); sdt.AppendChild(run); // Insert the content control at the beginning of the paragraph currentParagraph.InsertBefore(sdt, currentParagraph.FirstChild); } } return ReplaceAction.Skip; } }
解决方案
Aspose.Words默认的Range.Replace方法不支持跨段落/列表项的正则匹配,因为它的匹配逻辑基于单个文本节点或段落内的文本,无法跨节点组合匹配。要实现跨列表项的文本匹配并生成SDT,需要手动遍历文档节点,收集文本并进行匹配,再将匹配到的内容替换为SDT。
具体步骤如下:
- 遍历文档中的列表项段落,收集目标文本片段
- 检查是否存在跨段落的匹配(比如第一个列表项的"Hello"和第二个的"World")
- 找到匹配后,将两个列表项的内容合并到一个SDT中,移除原列表项并调整文档结构
以下是修改后的代码示例:
Document doc = new Document("YourDocumentPath.docx"); Regex regex = new Regex(@"Hello\s*World", RegexOptions.IgnoreCase); // 筛选出所有列表项段落 List<Paragraph> listItems = doc.GetChildNodes(NodeType.Paragraph, true) .Cast<Paragraph>() .Where(p => p.IsListItem) .ToList(); for (int i = 0; i < listItems.Count - 1; i++) { Paragraph firstItem = listItems[i]; Paragraph secondItem = listItems[i + 1]; // 拼接两个列表项的纯文本内容 string combinedText = $"{firstItem.ToString(SaveFormat.Text).Trim()}{secondItem.ToString(SaveFormat.Text).Trim()}"; if (regex.IsMatch(combinedText)) { // 创建块级SDT,用于容纳整个列表项内容 StructuredDocumentTag sdt = new StructuredDocumentTag(doc, SdtType.RichText, MarkupLevel.Block) { LockContentControl = false, LockContents = false, Title = "Clause", IsShowingPlaceholderText = false }; // 复制第一个列表项的内容到SDT sdt.AppendChild(firstItem.Clone(true)); // 复制第二个列表项的内容(跳过列表编号的特殊字符) foreach (Node node in secondItem.ChildNodes) { if (node.NodeType != NodeType.SpecialChar) sdt.AppendChild(node.Clone(true)); } // 将SDT插入原列表项位置 firstItem.ParentNode.InsertBefore(sdt, firstItem); // 删除原有的两个列表项 firstItem.Remove(); secondItem.Remove(); break; // 找到单个匹配后退出循环,若需批量匹配可移除此行 } } doc.Save("OutputDocument.docx");
代码说明
- 使用
MarkupLevel.Block创建块级SDT,确保能容纳多个段落内容 - 遍历列表项并拼接文本进行正则匹配,解决跨节点匹配的限制
- 复制原列表项内容到SDT时跳过列表编号的特殊字符,避免多余标记
- 替换完成后删除原列表项,保持文档结构整洁
如果需要支持更多段落的跨节点匹配,可以扩展逻辑,连续收集多个段落的文本进行匹配,再批量处理对应节点。
内容的提问来源于stack exchange,提问作者Suhas Parameshwara
相关产品推荐
相关产品推荐

