使用PowerShell提取匹配模式间指定范围的文本行
PowerShell提取文本文件中指定内容的脚本完善
需求规则
- 起始位置:找到连续两行(第一行是
LOAD SUMMARY,第二行是===========),提取起点为这两行的向上偏移2行的位置 - 结束位置:提取到包含
*****END LOAD SESSION*****的行的前一行为止
文本文件示例
maybe some more text some text SOme really important text LOAD SUMMARY. But I don't want to include this row. 023-03-06 13:57:55.719 <TASK_12836-WRITER_2_*_1> INFO: [WRT_8035] Load complete time: Mon Mar 06 13:57:55 2023 LOAD SUMMARY ============ WRT_8036 Target: MAPPING NAME (Instance Name: [MAPPING NAME]) WRT_8038 Inserted rows - Requested: 45147 Applied: 45147 Rejected: 0 Affected: 45147 2023-03-06 13:57:55.719 <TASK_12836-WRITER_2_*_1> INFO: [WRT_8043] *****END LOAD SESSION***** SOme more test LOAD SUMMARY > I don't want this row either
期望提取结果
023-03-06 13:57:55.719 <TASK_12836-WRITER_2_*_1> INFO: [WRT_8035] Load complete time: Mon Mar 06 13:57:55 2023 LOAD SUMMARY ============ WRT_8036 Target: MAPPING NAME (Instance Name: [MAPPING NAME]) WRT_8038 Inserted rows - Requested: 45147 Applied: 45147 Rejected: 0 Affected: 45147
当前编写的PowerShell脚本
# 指定要提取行的文件路径 $file = "C:\Users\YOURNAME\Desktop\Load Summary.txt" # 将文件内容读取为字符串数组 $content = Get-Content $file # 查找"LOAD SUMMARY"首次出现的行号 $startIndex = ($content | Select-String -Pattern "LOAD SUMMARY").LineNumber # 查找"*****END LOAD SESSION*****"首次出现的行号并减1 $endIndex = ($content | Select-String -Pattern "\*\*\*\*\*END LOAD SESSION\*\*\*\*\*").LineNumber - 1 # 在起始到结束索引范围内查找"============"行 for ($i = $startIndex; $i -le $endIndex; $i++) { if ($content[$i] -eq "============") { # 找到后将起始索引设为该行的下一行 $startIndex = $i + 1 break } } $startIndex = [int]$startIndex $endIndex = [int]$endIndex # 提取起始到结束索引间的行 $extractedLines = $content[$startIndex..$endIndex] # 输出提取的行 $extractedLines
完善后的PowerShell脚本
# 指定文件路径 $file = "C:\Users\YOURNAME\Desktop\Load Summary.txt" # 读取文件内容为数组(注意:Get-Content返回的数组索引从0开始,Select-String的LineNumber从1开始) $content = Get-Content $file # 遍历文件,找到连续的"LOAD SUMMARY"和"============"行,计算提取起始索引 $extractStartIndex = $content | ForEach-Object -Begin { $prevLine = $null } -Process { if ($prevLine -eq "LOAD SUMMARY" -and $_ -eq "============") { # 当前行是"============",向上偏移2行即当前索引减3(数组从0开始) return ($content.IndexOf($_) - 3) } $prevLine = $_ } | Select-Object -First 1 # 找到结束标记行的索引,提取到该行的前一行 $endMarker = $content | Select-String -Pattern "\*\*\*\*\*END LOAD SESSION\*\*\*\*\*" | Select-Object -First 1 $extractEndIndex = $endMarker.LineNumber - 2 # LineNumber转数组索引减1,再减1得到前一行索引 # 提取并输出结果 if ($extractStartIndex -ne $null -and $extractEndIndex -ne $null) { $content[$extractStartIndex..$extractEndIndex] } else { Write-Host "未找到符合条件的起始或结束标记" }
脚本说明
- 精准匹配连续的
LOAD SUMMARY和===========行,避免误匹配包含关键词的其他行 - 正确计算向上偏移2行的起始索引,符合需求
- 增加了异常判断,防止未找到标记时出错
- 统一处理数组索引与行号的转换逻辑,避免索引越界
内容的提问来源于stack exchange,提问作者Calico
相关产品推荐
相关产品推荐

