如何用C#提取字符串中指定词汇并统计其数量?
嘿,我帮你调整了代码,刚好能实现提取匹配词汇并统计数量的需求,一起来看看:
解决方案代码
using System; using System.Text.RegularExpressions; using System.Linq; public static class StringProcessor { public static void ExtractAndCountMatches(string input) { // 定义要匹配的目标词汇,把多词短语放在前面避免被拆分匹配 var targetWords = @"memory|buffer overflow|address space|stack overflow|call stack"; // 查找所有匹配项,忽略大小写 var matches = Regex.Matches(input, targetWords, RegexOptions.IgnoreCase); // 提取匹配内容、去重(忽略大小写)并整理成列表 var uniqueMatches = matches.Cast<Match>() .Select(m => m.Value.Trim()) .Distinct(StringComparer.OrdinalIgnoreCase) .ToList(); // 拼接成逗号分隔的字符串 var foundWords = string.Join(",", uniqueMatches); // 获取去重后的词汇数量 var count = uniqueMatches.Count; // 输出结果 Console.WriteLine($"Found: {foundWords}"); Console.WriteLine($"Count: {count}"); } } // 调用示例 var inputText = @"In software, a stack overflow occurs if the call stack pointer exceeds the stack bound. The call stack may consist of a limited amount of address space, often determined at the start of the program. The size of the call stack depends on many factors, including the programming language, machine architecture, multi-threading, and amount of available memory. When a program attempts to use more space than is available on the call stack (that is, when it attempts to access memory beyond the call stack's bounds, which is essentially a buffer overflow), the stack is said to overflow, typically resulting in a program crash."; StringProcessor.ExtractAndCountMatches(inputText);
代码说明
- 用
Regex.Matches替代原来的Replace,这样能捕获所有符合规则的匹配项,而不是直接替换掉 - 通过
Distinct(StringComparer.OrdinalIgnoreCase)实现忽略大小写的去重,确保像Stack Overflow和stack overflow被识别为同一个词汇 - 最后把去重后的词汇拼接成你要的格式,数量直接取去重列表的长度即可
运行结果
Found: stack overflow,call stack,address space,memory,buffer overflow
Count:5
内容的提问来源于stack exchange,提问作者Nat
相关产品推荐
相关产品推荐

