Python统计PowerShell导出日志,原文件返回空字典求原因
我是编程新手,最近遇到一个奇怪的问题:用PowerShell从远程服务器server01的Windows安全日志里过滤出ID为4663且Message不含name1、name2的记录,用Format-List和Out-File导出成文本文件后,用Python脚本统计用户名出现次数时,原文件返回空字典{},但把内容复制到新文件再运行脚本就得到了正确的统计结果{'name1':2,'name2':13,'name3':1,'name4':1,'name5':2,'name6':2}。两个文件看起来完全一样,字符位置也一致,想问问问题出在哪?
我的PowerShell代码
Get-WinEvent -LogName "Security" -ComputerName server01 | Where-Object {$_.ID -eq 4663} | where Message -CNotLike "*name1*" | where Message -CNotLike "*name2*" | Format-List -Property * | Out-File "C:\apowershell\winsec\events.txt"
我的Python代码
fhand = open('events2.txt') counts = dict() for line in fhand: if line.startswith(' Account Name:'): words = line.split() words.remove('Account') words.remove('Name:') for word in words: if word not in counts: counts[word] = 1 else: counts[word] += 1 print(counts)
日志记录样例
Subject: Security ID: S-1-5-21-495698755-754321212-623647154-4521 Account Name: name1 Account Domain: companydomain Logon ID: 0x8CB9C5024 Object: Object Server: Security Object Type: File Object Name: e:\share\file.txt Handle ID: 0x439c Resource Attributes: S:PAI Process Information: Process ID: 0x2de8 Process Name: C:\Windows\System32\memshell.exe Access Request Information: Accesses: Execute/Traverse Access Mask: 0x20
问题根源:编码不匹配
这问题我之前也碰到过好几次,核心原因是PowerShell的Out-File默认用UTF-16LE编码(也就是Unicode),但Python的open()函数默认用系统默认编码(通常是UTF-8或者GBK)打开文件,导致读取时无法正确解析文本内容,自然匹配不到' Account Name:'开头的行。
你复制到新文件的时候,系统的文本编辑器(比如记事本)会自动把UTF-16编码转成了UTF-8或者ANSI编码,这时候Python就能正常读取了。
解决办法有两种:
方法1:修改PowerShell导出时的编码
把Out-File改成指定UTF-8编码,这样Python默认就能正常读取:
Get-WinEvent -LogName "Security" -ComputerName server01 | Where-Object {$_.ID -eq 4663} | where Message -CNotLike "*name1*" | where Message -CNotLike "*name2*" | Format-List -Property * | Out-File "C:\apowershell\winsec\events.txt" -Encoding utf8
方法2:修改Python读取时的编码
在open()函数里指定UTF-16编码来匹配PowerShell的输出:
fhand = open('events.txt', encoding='utf-16') counts = dict() for line in fhand: if line.startswith(' Account Name:'): words = line.split() words.remove('Account') words.remove('Name:') for word in words: counts[word] = counts.get(word, 0) + 1 print(counts)
(顺便提一句,用counts.get(word, 0) + 1可以简化你的字典计数逻辑,不用写if-else判断~)
额外小建议
如果以后遇到类似“文件内容看起来一样但程序读不出来”的问题,首先要检查文件编码,可以用文本编辑器的“另存为”功能查看编码,或者用PowerShell的Get-Content -Path "events.txt" -Encoding Byte | Select-Object -First 4查看文件的BOM(UTF-16LE的BOM是FF FE,UTF-8的BOM是EF BB BF)。
内容的提问来源于stack exchange,提问作者durantejohn

