优化PowerShell处理大SFTP日志提取登录账号的效率
高效提取SFTP日志中成功登录账号的PowerShell方案
问题背景
有多份单文件约1GB、含数百万行的SFTP日志文件,需从中提取成功登录的账号。原脚本先筛选含“logged in”且排除“230”“530 Not logged in”的行导出为CSV,再拆分提取账号,但拆分阶段耗时极长、占用大量资源,希望找到更高效的实现方法。
原脚本
#Location of sftp logs in .txt $log_file = get-childitem -name C:\temp\logs_sftp\Logs #Extraction of lines where an account logged in foreach ($File in $log_file) { get-date Select-String -Path C:\Temp\logs_sftp\Logs\$file -Pattern "logged in" | select-string -Pattern "230" -NotMatch | select-string -Pattern "530 Not logged in" -NotMatch | Export-Csv C:\Temp\logs_sftp\Logs\$file.csv get-date } #Location of sftp logs with only the account lines in .csv #Split loop to extract only the account $log_file_csv = get-childitem -name C:\temp\logs_sftp\Logs\*.csv foreach ($File in $log_file_csv) { $ToSplit = Get-Content C:\temp\logs_mutualise\Logs\$file | Select-Object -Skip 2 $ToSplit | ForEach-Object { $aItems = $_ -split { $_ -eq " "} $aItems[$aItems.length-4] >> C:\Temp\logs_sftp\Result\export_split.csv } }
日志样本
"True","12","[5] Tue 01Nov22 00:00:00 - (63025700) User User_1 logged in","LogsServ-U01112022.txt","C:\Temp\logs_mutualise\Logs\LogsServ-U01112022.txt","logged in",,"System.Text.RegularExpressions.Match[]" "True","31","[5] Tue 01Nov22 00:00:00 - (63025701) User User_2 logged in","LogsServ-U01112022.txt","C:\Temp\logs_mutualise\Logs\LogsServ-U01112022.txt","logged in",,"System.Text.RegularExpressions.Match[]" "True","49","[5] Tue 01Nov22 00:00:00 - (63025702) User User_3 logged in","LogsServ-U01112022.txt","C:\Temp\logs_mutualise\Logs\LogsServ-U01112022.txt","logged in",,"System.Text.RegularExpressions.Match[]"
期望输出
User_1 User_2 User_3
优化方案
直接在筛选阶段提取账号,避免中间CSV文件的读写,同时用哈希表去重,最后批量写入结果,大幅提升效率:
# 定义日志目录和结果路径 $logDir = "C:\temp\logs_sftp\Logs" $resultPath = "C:\Temp\logs_sftp\Result\export_split.txt" # 用哈希表存储去重后的账号,键为账号名,值任意 $uniqueUsers = @{} # 遍历所有日志文件 Get-ChildItem -Path $logDir -Filter *.txt | ForEach-Object { Write-Host "处理文件: $($_.Name)" # 逐批量读取日志,匹配成功登录的行并提取账号 Get-Content -Path $_.FullName -ReadCount 1000 | ForEach-Object { foreach ($line in $_) { # 筛选包含logged in且排除230、530 Not logged in的行 if ($line -match "logged in" -and $line -notmatch "230" -and $line -notmatch "530 Not logged in") { # 用正则提取User后面的账号名 if ($line -match 'User (\w+?) logged in') { $user = $matches[1] $uniqueUsers[$user] = $true } } } } } # 将去重后的账号写入结果文件 $uniqueUsers.Keys | Out-File -Path $resultPath -Encoding utf8 Write-Host "处理完成,结果已保存至: $resultPath"
优化点说明
- 避免中间文件:直接在读取日志时提取账号,跳过原脚本中导出CSV再二次处理的步骤,减少磁盘I/O开销。
- 批量读取:使用
-ReadCount 1000批量读取行,比逐行读取更高效。 - 哈希表去重:利用哈希表的键唯一性自动去重,比后续再去重更节省资源。
- 正则精准提取:用正则表达式直接定位账号,比字符串拆分更可靠且高效。
- 最后批量写入:收集所有结果后一次性写入文件,避免多次追加写入的磁盘操作。
内容的提问来源于stack exchange,提问作者Matt
相关产品推荐
相关产品推荐

