如何使用Python将dir命令结果转换为CSV或JSON格式?
将Windows dir命令输出转换为CSV或JSON的Python实现方案
问题描述
我希望编写一个Python解析器,将
dir /a/s/od/ta命令的输出结果转换为CSV或JSON格式。例如,执行该命令得到的输出如下:纾佺鍗� C 涓殑纾佺娌掓湁妯欑堡銆�纾佺鍗�搴忚櫉: 123-1234 C:\ 鐨勭洰閷� 2009/07/14 涓婂崍 11:20 <DIR> PerfLogs 2009/07/14 涓嬪崍 01:08 <JUNCTION> Documents and Settings [C:\Users] 2018/03/14 涓嬪崍 04:09 12,796,198,912 hiberfil.sys 2018/03/14 涓嬪崍 04:09 17,061,601,280 pagefile.sys 2018/03/14 涓嬪崍 04:16 <DIR> Recovery 2018/03/14 涓嬪崍 04:17 <DIR> Users 2018/03/14 涓嬪崍 04:17 <DIR> $Recycle.Bin 2018/03/14 涓嬪崍 04:42 <DIR> Intel 2018/03/14 涓嬪崍 05:46 <DIR> Python27 2018/03/22 涓嬪崍 01:19 40 87E27E492B63 2018/03/31 涓婂崍 11:08 <DIR> py36 2018/05/04 涓嬪崍 06:12 <DIR> ProgramData 2018/05/04 涓嬪崍 06:12 <DIR> Program Files (x86) 2018/05/04 涓嬪崍 06:15 <DIR> Windows 2018/05/07 涓婂崍 10:17 <DIR> System Volume Information 2018/05/07 涓婂崍 10:19 <DIR> Config.Msi 2018/05/08 涓婂崍 10:39 <DIR> Program Files 3 鍊嬫獢妗� 29,857,800,232 浣嶅厓绲�期望转换后的格式(以表格形式示例):
============================================================ |**name |path |lastaccess |type** | |PerfLogs |C:\ |2009/07/14 涓婂崍 11:20 |<DIR> | |hiberfil.sys |C:\ |2018/03/14 涓嬪崍 04:09 |12,796,198,912| |pagefile.sys |C:\ |2018/03/14 涓嬪崍 04:09 |17,061,601,280| |... |... |... |... | ============================================================请问是否有适用的Python库可以实现该功能?如果没有相关库,能否提供实现建议?
解决方案
关于现成Python库
目前没有专门针对Windows dir命令输出解析的主流通用库,因为dir的输出格式高度依赖系统区域设置(比如语言、编码),格式并不固定,所以很少有库能覆盖所有场景。不过我们可以自己动手实现,步骤也不算复杂。
手动实现步骤建议
1. 正确获取dir命令的输出
首先要解决编码问题,Windows中文系统的dir输出通常使用gbk编码,用subprocess模块执行命令并指定编码:
import subprocess def get_dir_output(): # 执行dir命令,获取输出 result = subprocess.run( ["dir", "/a/s/od/ta"], capture_output=True, text=True, encoding="gbk" # 根据你的系统实际编码调整,比如utf-8 ) return result.stdout
2. 过滤有效行
dir的输出开头是卷信息,结尾是统计数据,需要过滤掉这些无效行:
- 跳过包含卷标识(如示例中的"鍗�")、目录说明("搴忚櫉")的开头行
- 跳过包含文件计数("鍊嬫獢妗�")、字节统计("浣嶅厓绲�")的结尾行
- 只保留包含日期格式(
YYYY/MM/DD)的有效行
3. 用正则解析每行数据
有效行分为两类:目录/链接行和文件行,用正则表达式匹配提取字段:
import re # 匹配目录/链接行:日期 时间 <类型> 名称 [目标路径] dir_pattern = re.compile(r"^(\d{4}/\d{2}/\d{2}) (\S+) (<[^>]+>) +(.+?)(?: \[(.+)\])?$") # 匹配文件行:日期 时间 文件大小 文件名 file_pattern = re.compile(r"^(\d{4}/\d{2}/\d{2}) (\S+) +([\d,]+) +(.+)$") def parse_line(line, root_path="C:\\"): line = line.strip() # 尝试匹配目录行 dir_match = dir_pattern.match(line) if dir_match: date, time, item_type, name, target = dir_match.groups() return { "name": name.strip(), "path": root_path, "lastaccess": f"{date} {time}", "type": item_type, "target": target.strip() if target else None # 可选,存储JUNCTION的目标路径 } # 尝试匹配文件行 file_match = file_pattern.match(line) if file_match: date, time, size, name = file_match.groups() return { "name": name.strip(), "path": root_path, "lastaccess": f"{date} {time}", "type": size.replace(",", "") # 可以转成整数,或保留原格式 } return None
4. 转换为CSV或JSON格式
- 转CSV:使用Python内置的
csv模块:import csv def write_to_csv(items, filename="dir_output.csv"): with open(filename, "w", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=["name", "path", "lastaccess", "type"]) writer.writeheader() writer.writerows(items) - 转JSON:使用Python内置的
json模块:import json def write_to_json(items, filename="dir_output.json"): with open(filename, "w", encoding="utf-8") as f: json.dump(items, f, ensure_ascii=False, indent=2)
5. 整合所有功能
把上面的函数组合起来,完成完整的解析流程:
def main(): dir_output = get_dir_output() lines = dir_output.split("\n") parsed_items = [] # 这里可以优化:从dir的开头行自动提取当前根路径,示例中暂时固定为C:\ current_path = "C:\\" for line in lines: item = parse_line(line, current_path) if item: parsed_items.append(item) # 输出为CSV和JSON write_to_csv(parsed_items) write_to_json(parsed_items) print(f"解析完成,共处理{len(parsed_items)}个项目") if __name__ == "__main__": main()
额外注意事项
- 编码适配:如果你的系统不是中文环境,需要调整
subprocess.run中的encoding参数(比如英文系统用utf-8) - 递归目录处理:如果使用
/s参数递归遍历子目录,dir会在输出中切换目录,需要解析输出中的目录切换行(如"XX目录的内容")来更新current_path - 特殊文件名:正则已经处理了带空格的文件名,但如果有更特殊的字符(如括号、中括号),可能需要微调正则表达式
内容的提问来源于stack exchange,提问作者BBBruce
相关产品推荐
相关产品推荐

