C#新手求助:如何拆分读取文本文件后得到的数组元素?
处理C#中文本文件人员信息拆分的方案
作为C#新手,咱们一步步来搞定这个问题——从读取文件内容,到拆分每个人员的独立信息块,再到提取具体字段,全程都是新手友好的操作~
第一步:读取文本文件内容
首先得把文件里的内容读进程序里,这里有两种常用方式,根据你的文件格式选就行:
- 如果文件里每行是完整的人员信息,用
File.ReadAllLines(@"文件路径")直接拿到字符串数组; - 如果人员信息是跨多行或者连续的块(像你给的例子那样),用
File.ReadAllText(@"文件路径")读取整个文本内容,后续处理更方便。
代码示例(带异常处理,避免新手踩坑):
using System.IO; // 必须引入这个命名空间才能操作文件 string fileContent = string.Empty; try { // 替换成你的实际文件路径,记得用@符号避免转义问题 fileContent = File.ReadAllText(@"C:\YourFiles\people_info.txt"); } catch (FileNotFoundException) { Console.WriteLine("文件没找到!请检查路径是否正确~"); } catch (IOException ex) { Console.WriteLine($"读取文件出错啦:{ex.Message}"); }
第二步:拆分单个人员的信息块
从你的示例内容来看,每个人员信息都带有ID #: XXXXXXX的标识,而且下一个人员的信息也会以ID #: 开头。咱们用正则表达式就能轻松把每个独立的人员信息块拆出来:
首先引入正则的命名空间:using System.Text.RegularExpressions;
然后写拆分代码:
// 正则逻辑:匹配从文本开头/上一个ID标识,到下一个ID标识/文本结尾的所有内容 // 非贪婪匹配能避免把多个人员信息合并成一个 MatchCollection personMatches = Regex.Matches( fileContent, @"(?<=^|ID #: ).*?(?=ID #: |$)", RegexOptions.Singleline ); // 把匹配到的信息转成列表,方便后续处理 List<string> personInfoList = new List<string>(); foreach (Match match in personMatches) { if (!string.IsNullOrWhiteSpace(match.Value)) { personInfoList.Add(match.Value.Trim()); } }
现在personInfoList里的每个元素,就是一个完整的人员信息字符串啦~
第三步:拆分每个人员的具体字段
接下来要把每个人员的信息拆成「姓名、ID、地址、全名、城市、州、邮编、SSN、就读项目」这些具体字段。咱们写一个专门的解析方法,用正则匹配固定格式的字段,新手也能看懂:
先定义一个存储人员信息的类(方便后续使用):
public class PersonInfo { public string ShortName { get; set; } // 比如示例里的Logan Babbleton public string ID { get; set; } public string Address { get; set; } public string FullName { get; set; } // 比如示例里的Mr. Logan M. Babbleton public string City { get; set; } public string State { get; set; } public string ZipCode { get; set; } public string SSN { get; set; } public string Program { get; set; } }
然后写解析方法:
public static PersonInfo ParseSinglePerson(string personText) { PersonInfo person = new PersonInfo(); // 提取ID(固定格式:ID #: 7位数字) Match idMatch = Regex.Match(personText, @"ID #: (\d{7})"); if (idMatch.Success) { person.ID = idMatch.Groups[1].Value; } // 提取SSN(固定格式:SSN: XXX-XX-XXXX) Match ssnMatch = Regex.Match(personText, @"SSN: (\w{3}-\w{2}-\w{4})"); if (ssnMatch.Success) { person.SSN = ssnMatch.Groups[1].Value; } // 提取就读项目(固定格式:当前就读项目:xxx) Match programMatch = Regex.Match(personText, @"当前就读项目:(.+?)(?=Mr\. |$)", RegexOptions.Singleline); if (programMatch.Success) { person.Program = programMatch.Groups[1].Value.Trim(); // 提取后把项目部分从文本里移除,方便处理剩余内容 personText = personText.Replace(programMatch.Value, "").Trim(); } // 提取全名(格式:Mr./Ms./Mrs. + 全名) Match fullNameMatch = Regex.Match(personText, @"(Mr\.|Ms\.|Mrs\.) (.+?)(?=\s+\w{2}\s+\d{5,6})"); if (fullNameMatch.Success) { person.FullName = $"{fullNameMatch.Groups[1].Value} {fullNameMatch.Groups[2].Value}".Trim(); personText = personText.Replace(fullNameMatch.Value, "").Trim(); } // 提取城市、州、邮编(州是2位大写字母,邮编是5-6位数字) Match locationMatch = Regex.Match(personText, @"(\w+(\s+\w+)*) (\w{2}) (\d{5,6})"); if (locationMatch.Success) { person.City = locationMatch.Groups[1].Value.Trim(); person.State = locationMatch.Groups[3].Value.Trim(); person.ZipCode = locationMatch.Groups[4].Value.Trim(); personText = personText.Replace(locationMatch.Value, "").Trim(); } // 剩下的内容就是短名称和地址,拆分一下 personText = personText.Replace($"ID #: {person.ID}", "").Trim(); int firstSpaceIndex = personText.IndexOf(' '); if (firstSpaceIndex != -1) { person.ShortName = personText.Substring(0, firstSpaceIndex).Trim(); person.Address = personText.Substring(firstSpaceIndex).Trim(); } return person; }
最后调用方法,把所有人员信息解析成对象列表:
List<PersonInfo> allPeople = new List<PersonInfo>(); foreach (string personText in personInfoList) { PersonInfo person = ParseSinglePerson(personText); allPeople.Add(person); } // 测试输出,看看结果对不对 foreach (var p in allPeople) { Console.WriteLine($"姓名:{p.ShortName} | ID:{p.ID} | 就读项目:{p.Program}"); }
新手注意事项
- 如果你的文本里有女性人员(用Ms./Mrs.),记得把正则里的
Mr\.改成(Mr\.|Ms\.|Mrs\.); - 如果邮编、ID的格式和示例不一样,要对应调整正则里的数字位数;
- 所有用到的命名空间(
System.IO、System.Text.RegularExpressions、System.Collections.Generic)都要在文件顶部引入哦。
内容的提问来源于stack exchange,提问作者rcarcia02
相关产品推荐
相关产品推荐

