You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python Polars将LDIF文件转换为表格型DataFrame?

使用Python Polars将LDIF文件转换为表格格式

问题背景

我有多个如下格式的LDIF文件:

dn: uid=jdoe,ou=People,dc=example,dc=com
changetype: add
objectClass: inetOrgPerson
uid: jdoe
cn: John Doe
sn: Doe
mail: jdoe@example.com

dn: uid=asmith,ou=People,dc=example,dc=com
changetype: add
objectClass: inetOrgPerson
uid: asmith
cn: Alice Smith
sn: Smith
mail: asmith@example.com

在R中我常用unnest()函数将这类数据规范化为表格格式,目标输出如下:

dnchangetypeobjectClassuidcnsnmail
uid=jdoe,ou=People,dc=example,dc=comaddinetOrgPersonjdoeJohn DoeDoejdoe@example.com
uid=asmith,ou=People,dc=example,dc=comaddinetOrgPersonasmithAlice SmithSmithasmith@example.com

请问如何用Python Polars实现这一转换?


解决方案

核心思路是按空行分割LDIF中的独立条目,解析每个条目的键值对,再通过Polars转换为标准表格,具体步骤如下:

1. 处理单个LDIF文件

import polars as pl

# 读取文件并过滤无效行(空行、注释行)
with open("example.ldif", "r") as f:
    lines = [line.strip() for line in f if line.strip() and not line.startswith("#")]

# 按dn:分割为独立条目组
entry_groups = []
current_entry = []
for line in lines:
    if line.startswith("dn:"):
        if current_entry:
            entry_groups.append(current_entry)
        current_entry = [line]
    else:
        current_entry.append(line)
if current_entry:
    entry_groups.append(current_entry)

# 解析每个条目为字典,再转为Polars DataFrame
parsed_entries = []
for entry in entry_groups:
    entry_dict = {}
    for line in entry:
        key, value = line.split(":", 1)
        entry_dict[key.strip()] = value.strip()
    parsed_entries.append(entry_dict)

df = pl.DataFrame(parsed_entries)
print(df)

2. 批量处理多个LDIF文件

如果有多个LDIF文件,可以通过循环读取合并结果:

import polars as pl
import glob

all_entries = []

# 遍历所有ldif文件
for ldif_path in glob.glob("*.ldif"):
    with open(ldif_path, "r") as f:
        lines = [line.strip() for line in f if line.strip() and not line.startswith("#")]
    
    current_entry = []
    for line in lines:
        if line.startswith("dn:"):
            if current_entry:
                # 解析当前条目为字典
                entry_dict = {k.strip(): v.strip() for k, v in (l.split(":", 1) for l in current_entry)}
                all_entries.append(entry_dict)
            current_entry = [line]
        else:
            current_entry.append(line)
    # 处理文件最后一个条目
    if current_entry:
        entry_dict = {k.strip(): v.strip() for k, v in (l.split(":", 1) for l in current_entry)}
        all_entries.append(entry_dict)

# 生成最终表格
final_df = pl.DataFrame(all_entries)
print(final_df)

说明

  • 代码会自动处理条目间的字段差异,缺失字段会被填充为null
  • 支持过滤LDIF文件中的注释行(以#开头的行)

内容的提问来源于stack exchange,提问作者CT_369

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 13:43:15