C#中如何使用Regex提取字符串中的多个指定部分?
C# 用正则提取指定格式化的字符串内容
针对你需要提取被<em><strong></strong></em>包裹内容的需求,用正则表达式可以一次性完成提取,比多次string.Split()高效得多。
正则表达式方案
可以使用带正向/反向断言的正则,精准匹配被指定标签包裹的内容:
var pattern = @"(?<=<em><strong>)([A-Za-z]{3}|[0-9]+)(?=</strong></em>)";
(?<=<em><strong>):反向断言,确保目标内容前是指定起始标签([A-Za-z]{3}|[0-9]+):捕获组,匹配两种规则:要么是固定3个字母(对应你说的第一个固定长度部分),要么是任意长度数字(对应后面的数字部分)(?=</strong></em>):正向断言,确保目标内容后是指定结束标签
C# 代码示例
using System; using System.Text.RegularExpressions; class Program { static void Main() { string input = "TCS-<em><strong>TST</strong></em>.MSL-M365-SPO.S<em><strong>8629</strong></em>-O<em><strong>2887</strong></em>.Engagement"; var pattern = @"(?<=<em><strong>)([A-Za-z]{3}|[0-9]+)(?=</strong></em>)"; MatchCollection matches = Regex.Matches(input, pattern); foreach (Match match in matches) { Console.WriteLine(match.Value); } // 输出结果: // TST // 8629 // 2887 } }
进阶优化(精准匹配固定格式)
如果你需要明确区分不同位置的内容,也可以直接匹配整个字符串的固定结构,直接提取对应位置的目标值:
var fullPattern = @"TCS-<em><strong>([A-Za-z]{3})</strong></em>\.MSL-M365-SPO\.S<em><strong>(\d+)</strong></em>-O<em><strong>(\d+)</strong></em>\.Engagement"; Match fullMatch = Regex.Match(input, fullPattern); if (fullMatch.Success) { string firstPart = fullMatch.Groups[1].Value; // TST string secondPart = fullMatch.Groups[2].Value; // 8629 string thirdPart = fullMatch.Groups[3].Value; // 2887 }
这种方式完全贴合你给出的字符串格式,能避免匹配到无关的标签内容,提取结果更精准。
内容的提问来源于stack exchange,提问作者Sachin
相关产品推荐
相关产品推荐

