求PowerShell脚本:将文本文件中不含冒号的行合并至前一行
求PowerShell脚本:将文本文件中不含冒号的行合并至前一行
兄弟我太懂你这种被多行数据打乱格式的糟心情况了!之前逐行处理没上下文确实卡壳,给你整了个高效的PowerShell方案,不用反复折腾行号,大文件处理也稳得一批。
核心思路就是线性遍历一次所有行,用一个变量追踪当前正在构建的条目行:遇到带冒号的新条目时,先输出之前构建好的行,再开始新的条目;遇到不带冒号的行,直接把它拼接到当前条目行后面,用::做分隔。
基础版脚本(适合大多数文件)
# 替换成你的输入文件路径 $inputFile = "C:\your\input\file.txt" # 替换成你的输出文件路径 $outputFile = "C:\your\output\merged-file.txt" # 读取所有行到数组 $allLines = Get-Content -Path $inputFile $currentEntry = $null foreach ($line in $allLines) { # 跳过空行(如果不需要跳过,直接删掉这段判断) if ([string]::IsNullOrWhiteSpace($line)) { continue } if ($line -match ':') { # 如果已经有正在构建的条目,先输出到文件 if ($currentEntry) { $currentEntry | Out-File -Path $outputFile -Append -Encoding utf8 } # 开始新的条目行 $currentEntry = $line } else { # 将当前行拼接到上一条目,用::分隔 $currentEntry += " :: $line" } } # 输出最后一个条目(循环结束后最后一行还没输出) if ($currentEntry) { $currentEntry | Out-File -Path $outputFile -Append -Encoding utf8 }
超大文件优化版(流式读取,低内存占用)
如果你的文件是几GB级别的超大文件,用上面的Get-Content会把整个文件加载到内存,可能会卡。改用.NET的ReadLines流式读取,边读边处理,内存占用极低:
$inputFile = "C:\your\large-input\file.txt" $outputFile = "C:\your\output\merged-file.txt" # 先清空输出文件(如果需要覆盖原有内容) if (Test-Path -Path $outputFile) { Clear-Content -Path $outputFile } $currentEntry = $null # 流式读取每一行,不一次性加载全部内容 foreach ($line in [System.IO.File]::ReadLines($inputFile)) { if ([string]::IsNullOrWhiteSpace($line)) { continue } if ($line -match ':') { if ($currentEntry) { Add-Content -Path $outputFile -Value $currentEntry -Encoding utf8 } $currentEntry = $line } else { $currentEntry += " :: $line" } } # 输出最后一个条目 if ($currentEntry) { Add-Content -Path $outputFile -Value $currentEntry -Encoding utf8 }
效果验证
拿你给的示例输入测试:
ENTRY: XYZ COMMENT: This is a comment that spans over multiple lines just to make life difficult ENTRY: 123
处理后输出 exactly 你想要的结果:
ENTRY: XYZ COMMENT: This is a comment :: that spans over multiple lines :: just to make life difficult ENTRY: 123
自定义调整小技巧
- 如果不想跳过空行,直接删掉
if ([string]::IsNullOrWhiteSpace($line)) { continue }这段代码 - 分隔符
::可以改成你喜欢的样式,比如|或者直接空格 - 编码格式
Encoding utf8可以根据你的文件实际情况调整,比如改成Encoding default或者Encoding utf8BOM
备注:内容来源于stack exchange,提问作者seagull
相关产品推荐
相关产品推荐

