C#正则表达式提取文本文件中最后一个数据表格
提取文本文件中最后一个特定格式的数据表格
嘿,我懂你要解决的问题了——你的文本文件里重复出现好几次那种My data Statistics:开头的统计表格,现在你已经能用正则找出所有匹配的开头,但想要抓最后那个完整的表格对吧?
这里有两种实用的方法可以帮你搞定:
方法1:基于你现有代码的改进
你已经通过Regex.Matches拿到了所有表格开头的匹配项,那直接取最后一个匹配的起始位置,然后截取从该位置到文本末尾的内容就行,这样就能得到最后一个完整表格:
using System.IO; using System.Text.RegularExpressions; // 读取文件内容 StreamReader reader = new StreamReader("你的文件路径.txt"); string text = reader.ReadToEnd(); reader.Close(); // 找到所有表格开头的匹配 MatchCollection matches = Regex.Matches(text, "My data Statistics:"); if (matches.Count > 0) { // 获取最后一个匹配项的起始索引 int lastTableStartIndex = matches[matches.Count - 1].Index; // 截取从该索引到文本末尾的内容,就是最后一个表格 string lastTableContent = text.Substring(lastTableStartIndex); // 这里可以添加对表格内容的处理逻辑,比如解析字段 Console.WriteLine(lastTableContent); } else { Console.WriteLine("未找到符合格式的统计表格"); }
方法2:用正则直接匹配最后一个表格
如果你想更简洁,也可以直接写一个正则表达式,一次性匹配出最后一个完整的表格。这个正则会匹配My data Statistics:开头,然后匹配所有内容直到后面再也没有新的表格开头为止:
using System.IO; using System.Text.RegularExpressions; StreamReader reader = new StreamReader("你的文件路径.txt"); string text = reader.ReadToEnd(); reader.Close(); // 正则解释:(?s)让.匹配换行符,.*匹配所有内容,(?!...)确保后面没有新的表格开头 string regexPattern = @"My data Statistics:(?s).*(?!My data Statistics:)"; Match lastTableMatch = Regex.Match(text, regexPattern); if (lastTableMatch.Success) { string lastTableContent = lastTableMatch.Value; Console.WriteLine(lastTableContent); } else { Console.WriteLine("未找到符合格式的统计表格"); }
小提示
如果你的文件特别大,ReadToEnd()可能会占用较多内存,这种情况下可以考虑逐行读取并记录最后一次出现表格开头后的所有行,但对于常规大小的文件,上面两种方法足够高效好用。
内容的提问来源于stack exchange,提问作者Erofh Tor
相关产品推荐
相关产品推荐

