C#中TXT文件转可检索结构化数据及字段拆分求助
Hey there! I totally get how messy unstructured text files can make your brain feel—let's break this down into simple, actionable steps to get your data organized properly. Here's a straightforward approach that handles multi-line values for Field3/Field4, skips useless fields, and keeps everything linked to Field1 as you need.
Step 1: Define a Data Model
First, create a class to represent each "group" of data tied to a single Field1. This will let you easily access all related Field2, Field3, and Field4 entries:
public class DataRecord { public string Field1 { get; set; } public string Field2 { get; set; } public List<string> Field3Entries { get; set; } = new List<string>(); public List<string> Field4Entries { get; set; } = new List<string>(); }
Field1acts as your key identifier for each data group.Field3EntriesandField4Entriesare lists to hold multiple entries (including multi-line values) linked to the parent Field1.
Step 2: Parse the TXT File
This code reads through your file line by line, handles multi-line values, skips Useless fields, and builds your list of DataRecord objects:
using System; using System.Collections.Generic; using System.IO; using System.Text; public class TxtParser { public static List<DataRecord> ParseTxtFile(string filePath) { var records = new List<DataRecord>(); DataRecord currentRecord = null; string activeMultiLineField = null; // Tracks if we're collecting multi-line data for Field3/4 var multiLineValueBuilder = new StringBuilder(); foreach (var rawLine in File.ReadLines(filePath)) { var line = rawLine.Trim(); if (string.IsNullOrWhiteSpace(line)) continue; // Check if this line starts a new field (contains ": ") if (line.Contains(": ")) { // Save any unfinished multi-line value first if (activeMultiLineField != null && currentRecord != null) { SaveMultiLineValue(currentRecord, activeMultiLineField, multiLineValueBuilder); activeMultiLineField = null; multiLineValueBuilder.Clear(); } // Split into field name and value var splitPos = line.IndexOf(": "); var fieldName = line.Substring(0, splitPos).Trim(); var fieldValue = line.Substring(splitPos + 2).Trim(); // Skip all Useless fields immediately if (fieldName.StartsWith("Useless")) continue; // Handle each valid field switch (fieldName) { case "Field1": // Start a new record when we hit a new Field1 if (currentRecord != null) records.Add(currentRecord); currentRecord = new DataRecord { Field1 = fieldValue }; break; case "Field2": currentRecord?.Field2 = fieldValue; break; case "Field3": // Start collecting multi-line data for Field3 activeMultiLineField = "Field3"; multiLineValueBuilder.Append(fieldValue); break; case "Field4": // Start collecting multi-line data for Field4 activeMultiLineField = "Field4"; multiLineValueBuilder.Append(fieldValue); break; } } else { // This line is part of a multi-line Field3/4 value if (activeMultiLineField != null && currentRecord != null) { multiLineValueBuilder.AppendLine(line); } // Ignore unrecognized lines that aren't part of a multi-line field } } // Clean up: save any remaining multi-line value and add the last record if (activeMultiLineField != null && currentRecord != null) { SaveMultiLineValue(currentRecord, activeMultiLineField, multiLineValueBuilder); } if (currentRecord != null) { records.Add(currentRecord); } return records; } private static void SaveMultiLineValue(DataRecord record, string fieldName, StringBuilder valueBuilder) { var finalValue = valueBuilder.ToString().Trim(); switch (fieldName) { case "Field3": record.Field3Entries.Add(finalValue); break; case "Field4": record.Field4Entries.Add(finalValue); break; } } }
Step 3: Query Your Data
Once you've parsed the file, retrieving associated data is a breeze with LINQ. For example:
// Parse the file var allRecords = TxtParser.ParseTxtFile("your_data.txt"); // Find all Field3 entries linked to a specific Field1 value string targetField1 = "YourTargetData"; var matchingRecord = allRecords.FirstOrDefault(r => r.Field1 == targetField1); if (matchingRecord != null) { Console.WriteLine($"Field3 entries for {targetField1}:"); foreach (var field3 in matchingRecord.Field3Entries) { Console.WriteLine($"- {field3}"); } } // Or get all Field4 entries across all records var allField4Entries = allRecords.SelectMany(r => r.Field4Entries);
Key Notes:
- Multi-line Handling: If Field3/4 values span multiple lines (without the
Field3:prefix on subsequent lines), the code will append those lines to the current value. - Useless Fields: Any field starting with "Useless" is skipped entirely—no processing, no storage.
- Edge Cases: If the file starts with a Field3/4 (no preceding Field1), those entries will be ignored since they can't be linked to a key. You can adjust this logic if needed.
内容的提问来源于stack exchange,提问作者Mrcl1450

