C#读取TXT文件换行异常:单词统计结果不符求解决方案
解决TXT文件换行导致单词统计错误的问题
嘿,我来帮你搞定这个换行引发的统计bug!你的核心问题是分割文本时没把换行符(\n)和回车符(\r)加入分隔符列表,导致换行前后的单词被当成了奇怪的组合(比如你看到的world coo),自然统计结果就不对了。下面给你两种靠谱的解决方案:
方案1:改进Split方法,覆盖换行分隔符
直接把换行、回车符加到分隔符数组里,再加上StringSplitOptions.RemoveEmptyEntries避免空字符串干扰:
string text = File.ReadAllText(@"C:\Users\derka\Documents\test.txt"); // 补充换行符、回车符到分隔符集合 char[] separators = new char[] { ' ', ',', '.', ':', '\t', '\n', '\r' }; string[] array = text.Split(separators, StringSplitOptions.RemoveEmptyEntries); // 接下来统计单词次数 Dictionary<string, int> wordCount = new Dictionary<string, int>(); foreach (string word in array) { // 可选:统一转小写,避免Hello和hello被当成不同单词 string lowerWord = word.ToLower(); if (wordCount.ContainsKey(lowerWord)) { wordCount[lowerWord]++; } else { wordCount[lowerWord] = 1; } }
方案2:用正则表达式更优雅地提取单词
正则能一次性匹配所有单词(自动忽略标点、空格、换行等非单词字符),完美解决各种分隔符问题。你说不知道怎么用foreach处理,看这里:
using System.Text.RegularExpressions; string text = File.ReadAllText(@"C:\Users\derka\Documents\test.txt"); // 匹配所有单词(\w+ 表示匹配一个或多个字母/数字/下划线) MatchCollection wordMatches = Regex.Matches(text, @"\w+"); Dictionary<string, int> wordCount = new Dictionary<string, int>(); foreach (Match match in wordMatches) { string word = match.Value.ToLower(); if (wordCount.ContainsKey(word)) { wordCount[word]++; } else { wordCount[word] = 1; } }
最后输出结果到文件
统计完成后,把结果写入新的TXT文件:
using (StreamWriter writer = new StreamWriter(@"C:\Users\derka\Documents\word_count_result.txt")) { foreach (var kvp in wordCount) { writer.WriteLine($"{kvp.Key}, {kvp.Value}"); } }
这样处理后,你的测试文件里的world就能正确统计出3次啦!
内容的提问来源于stack exchange,提问作者dasdzoni
相关产品推荐
相关产品推荐

