如何使用Python Polars将LDIF文件转换为表格型DataFrame?
使用Python Polars将LDIF文件转换为表格格式
问题背景
我有多个如下格式的LDIF文件:
dn: uid=jdoe,ou=People,dc=example,dc=com changetype: add objectClass: inetOrgPerson uid: jdoe cn: John Doe sn: Doe mail: jdoe@example.com dn: uid=asmith,ou=People,dc=example,dc=com changetype: add objectClass: inetOrgPerson uid: asmith cn: Alice Smith sn: Smith mail: asmith@example.com
在R中我常用unnest()函数将这类数据规范化为表格格式,目标输出如下:
| dn | changetype | objectClass | uid | cn | sn | |
|---|---|---|---|---|---|---|
| uid=jdoe,ou=People,dc=example,dc=com | add | inetOrgPerson | jdoe | John Doe | Doe | jdoe@example.com |
| uid=asmith,ou=People,dc=example,dc=com | add | inetOrgPerson | asmith | Alice Smith | Smith | asmith@example.com |
请问如何用Python Polars实现这一转换?
解决方案
核心思路是按空行分割LDIF中的独立条目,解析每个条目的键值对,再通过Polars转换为标准表格,具体步骤如下:
1. 处理单个LDIF文件
import polars as pl # 读取文件并过滤无效行(空行、注释行) with open("example.ldif", "r") as f: lines = [line.strip() for line in f if line.strip() and not line.startswith("#")] # 按dn:分割为独立条目组 entry_groups = [] current_entry = [] for line in lines: if line.startswith("dn:"): if current_entry: entry_groups.append(current_entry) current_entry = [line] else: current_entry.append(line) if current_entry: entry_groups.append(current_entry) # 解析每个条目为字典,再转为Polars DataFrame parsed_entries = [] for entry in entry_groups: entry_dict = {} for line in entry: key, value = line.split(":", 1) entry_dict[key.strip()] = value.strip() parsed_entries.append(entry_dict) df = pl.DataFrame(parsed_entries) print(df)
2. 批量处理多个LDIF文件
如果有多个LDIF文件,可以通过循环读取合并结果:
import polars as pl import glob all_entries = [] # 遍历所有ldif文件 for ldif_path in glob.glob("*.ldif"): with open(ldif_path, "r") as f: lines = [line.strip() for line in f if line.strip() and not line.startswith("#")] current_entry = [] for line in lines: if line.startswith("dn:"): if current_entry: # 解析当前条目为字典 entry_dict = {k.strip(): v.strip() for k, v in (l.split(":", 1) for l in current_entry)} all_entries.append(entry_dict) current_entry = [line] else: current_entry.append(line) # 处理文件最后一个条目 if current_entry: entry_dict = {k.strip(): v.strip() for k, v in (l.split(":", 1) for l in current_entry)} all_entries.append(entry_dict) # 生成最终表格 final_df = pl.DataFrame(all_entries) print(final_df)
说明
- 代码会自动处理条目间的字段差异,缺失字段会被填充为
null - 支持过滤LDIF文件中的注释行(以
#开头的行)
内容的提问来源于stack exchange,提问作者CT_369
相关产品推荐
相关产品推荐

