You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在C#中处理特殊格式多行文本并生成指定输出的技术需求

C#处理特殊格式多行文本并格式化输出方案

没问题,我来帮你搞定这个C#里处理混乱多行文本并格式化输出的需求!首先先明确你的输入和目标:

原始文本内容如下:
||Names : XYZ DJ Age : 23 Years Location: New York; end;'
2018-03-20 11:59:59.397, mnx=0x0000700, pid=90c9ac, xSG: dlgID:34 AppDlg:774 params:
2018-03-20 11:59:59.397, mnx=0x700000, pid=090c9ac, lBG: OPCDManager::Response: 0x7f083
2018-03-20 11:59:59.397, mxn=0x000070, pid=f90c9ac, lBG: DlgID:37774 sess:'990' conID:1 dlClose:false params:

你的目标是提取关键信息后,按XYZ DJ-23 Years-New York-...的格式拼接并写入文件。下面是具体的实现方案:

核心思路

我们分三部分处理文本:

  1. 解析第一行的用户基础信息(姓名、年龄、地点)
  2. 提取后续日志行的关键字段(时间、进程ID、核心日志内容)
  3. 用-作为分隔符拼接所有信息,最后写入文件

完整C#代码实现

using System;
using System.IO;
using System.Text;
using System.Text.RegularExpressions;

class LogFormatter
{
    static void Main()
    {
        // 替换成你的原始多行文本(也可以从文件直接读取)
        string rawLogText = @"||Names : XYZ DJ Age : 23 Years Location: New York; end;' 
2018-03-20 11:59:59.397, mnx=0x0000700, pid=90c9ac, xSG: dlgID:34 AppDlg:774 params: 
2018-03-20 11:59:59.397, mnx=0x700000, pid=090c9ac, lBG: OPCDManager::Response: 0x7f083 
2018-03-20 11:59:59.397, mxn=0x000070, pid=f90c9ac, lBG: DlgID:37774 sess:'990' conID:1 dlClose:false params:";

        // 分割文本为独立行,过滤空行
        string[] logLines = rawLogText.Split(new[] { Environment.NewLine }, StringSplitOptions.RemoveEmptyEntries);

        StringBuilder formattedOutput = new StringBuilder();

        // 第一步:处理第一行的用户信息
        if (logLines.Length > 0)
        {
            var userInfoMatch = Regex.Match(logLines[0], @"\|{2}Names\s*:\s*(.*?)\s*Age\s*:\s*(.*?)\s*Location:\s*(.*?);");
            if (userInfoMatch.Success)
            {
                string userName = userInfoMatch.Groups[1].Value.Trim();
                string userAge = userInfoMatch.Groups[2].Value.Trim();
                string userLocation = userInfoMatch.Groups[3].Value.Trim();
                
                formattedOutput.Append($"{userName}-{userAge}-{userLocation}");
            }
        }

        // 第二步:处理后续日志行
        for (int i = 1; i < logLines.Length; i++)
        {
            string currentLine = logLines[i].Trim();
            
            // 提取日志时间(开头的标准时间格式)
            var timeMatch = Regex.Match(currentLine, @"^(\d{4}-\d{2}-\d{2}\s\d{2}:\d{2}:\d{2}\.\d{3})");
            // 提取进程ID pid字段
            var pidMatch = Regex.Match(currentLine, @"pid=([0-9a-fA-F]+)");
            // 提取日志核心内容(逗号后的部分)
            var contentMatch = Regex.Match(currentLine, @",\s*(.*)$");

            // 处理匹配失败的情况,给默认值避免空内容
            string logTime = timeMatch.Success ? timeMatch.Groups[1].Value : "Unknown_Time";
            string processId = pidMatch.Success ? pidMatch.Groups[1].Value : "Unknown_PID";
            string logContent = contentMatch.Success ? contentMatch.Groups[1].Value.Trim() : "Unknown_Content";

            // 按格式拼接到结果中
            formattedOutput.Append($"-[{logTime}][PID:{processId}]{logContent}");
        }

        // 第三步:写入到目标文本文件
        string outputFilePath = @"C:\Your_Target_Path\formatted_log.txt";
        // 如果需要相对路径,可以用 @"formatted_log.txt" 会生成在程序运行目录下
        File.WriteAllText(outputFilePath, formattedOutput.ToString());

        Console.WriteLine("文本处理完成!格式化后的内容已写入指定文件。");
    }
}

关键细节说明

  • 正则匹配的灵活性:正则表达式里用了\s*来匹配任意数量的空格,即使原始文本里字段前后空格不一致也能正常提取;如果字段名有大小写变化,可以给正则加上RegexOptions.IgnoreCase参数。
  • 容错处理:每个字段匹配都做了失败判断,给了默认值,避免因为某一行格式异常导致整个程序崩溃。
  • 可扩展性:如果需要调整输出格式,比如把[...]换成其他分隔符,或者新增提取字段,只需要修改正则和拼接逻辑即可。

内容的提问来源于stack exchange,提问作者adsf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:54:36