如何用Split(无Substring)、Regex拆分含编号字符串并优化性能?
字符串编号与标题内容拆分的三种C#实现方案
需求说明
现有一批字符串,部分包含十进制数字或字母组成的多级编号(如1.、2.1.、A.等),部分无编号。需将含编号的字符串拆分为编号(如1、2.1、A)与标题内容两部分,无编号的字符串则保留原内容。
示例输入
1. Title1 Contains Space and Characters2.1. Title2.1 Contains Space and CharactersA. Title A Contains Space and CharactersTitle which does not contains any Display Code
预期输出
提取编号和对应标题内容,无编号的字符串编号字段为空,标题保留原文。
一、仅使用Split方法实现(不依赖Substring)
完全通过Split方法完成拆分,避免使用Substring,核心思路是先按空格拆分字符串,再判断第一部分是否为合法编号:
string input = "2.1. Title2.1 Contains Space and Characters"; string displayCode = string.Empty; string title = input; // 按空格拆分,最多拆分为2个部分 var splitParts = input.Split(new[] { ' ' }, 2); if (splitParts.Length > 1) { var candidateCode = splitParts[0]; // 验证候选编号:以.结尾,且仅包含字母、数字和. if (candidateCode.EndsWith(".") && candidateCode.All(c => char.IsLetterOrDigit(c) || c == '.')) { displayCode = candidateCode.TrimEnd('.'); // 去除标题前的多余空格 title = splitParts[1].TrimStart(); } } Console.WriteLine($"Display Code: {displayCode} Title: {title}");
二、使用正则表达式实现
通过正则表达式精准匹配编号格式,支持多级编号(如2.1),同时处理编号后的多个空格:
using System.Text.RegularExpressions; string input = "A. Title A Contains Space and Characters"; string displayCode = string.Empty; string title = input; // 正则表达式:匹配开头的合法编号,以及后续的标题内容 // 分组1:编号(允许多级,如2.1);分组2:标题内容 var regexPattern = @"^([A-Za-z0-9]+(?:\.[A-Za-z0-9]+)*)\.\s+(.*)$"; var regex = new Regex(regexPattern, RegexOptions.Compiled); var matchResult = regex.Match(input); if (matchResult.Success) { displayCode = matchResult.Groups[1].Value; title = matchResult.Groups[2].Value; } Console.WriteLine($"Display Code: {displayCode} Title: {title}");
说明:添加
RegexOptions.Compiled可预编译正则表达式,提升多次匹配时的性能。
三、最优高效实现方式
若需处理大量字符串,优先选择直接定位字符位置的方案,避免Split或正则表达式带来的额外内存分配与开销。核心思路是通过IndexOf快速定位.和空格的位置,验证编号合法性后拆分:
public static (string DisplayCode, string Title) SplitTitle(string input) { if (string.IsNullOrWhiteSpace(input)) return (string.Empty, input); // 先尝试匹配". "(点+单个空格)的情况 int dotSpaceIndex = input.IndexOf(". "); if (dotSpaceIndex != -1) { string candidateCode = input.Substring(0, dotSpaceIndex); // 验证编号仅包含字母、数字和. if (candidateCode.All(c => char.IsLetterOrDigit(c) || c == '.')) { string title = input.Substring(dotSpaceIndex + 2).TrimStart(); return (candidateCode, title); } } // 处理编号后多个空格的情况(如A. Title) int dotIndex = input.IndexOf('.'); if (dotIndex != -1 && dotIndex + 1 < input.Length) { int spaceStartIndex = dotIndex + 1; // 跳过所有连续空格 while (spaceStartIndex < input.Length && char.IsWhiteSpace(input[spaceStartIndex])) { spaceStartIndex++; } // 确认存在有效标题内容 if (spaceStartIndex > dotIndex + 1 && spaceStartIndex < input.Length) { string candidateCode = input.Substring(0, dotIndex); if (candidateCode.All(c => char.IsLetterOrDigit(c) || c == '.')) { string title = input.Substring(spaceStartIndex); return (candidateCode, title); } } } // 无合法编号,返回原内容 return (string.Empty, input); } // 使用示例 var result1 = SplitTitle("1. Title1 Contains Space and Characters"); Console.WriteLine($"Display Code: {result1.DisplayCode} Title: {result1.Title}"); var result2 = SplitTitle("Title which does not contains any Display Code"); Console.WriteLine($"Display Code: {result2.DisplayCode} Title: {result2.Title}");
说明:该方案通过直接操作字符索引,减少了数组创建(Split)和正则匹配的开销,在处理海量字符串时性能最优。
内容的提问来源于stack exchange,提问作者PavanKumar GVVS
相关产品推荐
相关产品推荐

