You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Polars是否有类似Pandas read_fwf的定宽文件读取函数?

Python Polars 读取固定宽度格式文件的实现方法

Polars 目前没有像 Pandas read_fwf 那样的专属固定宽度文件读取函数,但可以通过逐行读取+按位置截取的方式轻松实现需求,完全适配你提供的位置式schema。

实现步骤与代码示例

  1. 定义列位置schema:按需求维护列名与对应的起始位置、截取长度(示例中schema为1-based索引,后续需转换为Python的0-based切片)
  2. 读取文本行:将文件读取为文本行集合,过滤空行
  3. 按位置提取列:遍历schema,用字符串切片提取各列内容
  4. (可选)指定数据类型:对提取的列进行类型转换
import polars as pl

# 你的列位置schema(1-based起始位置,截取长度)
schema_dict = {
    "Postal_codeOM": (1, 6),
    "FSA": (7, 3),
    # 补充其他列的位置定义...
}

# 读取文件并处理空行
df = pl.read_text("fixed_width.txt") \
       .str.split("\n") \
       .explode() \
       .filter(pl.col("text").str.lengths() > 0)

# 按schema提取所有列
for col_name, (start_pos, length) in schema_dict.items():
    # 转换为0-based切片:起始位置减1,截取指定长度
    df = df.with_columns(
        pl.col("text").str.slice(start_pos - 1, length).alias(col_name)
    )

# 移除原始文本列,得到最终结构化数据
df = df.drop("text")
print(df)

大文件优化方案

如果处理超大文件,建议用流式扫描的方式避免内存过载:

# 流式处理大文件
scan = pl.scan_text("fixed_width.txt") \
         .str.split("\n") \
         .explode() \
         .filter(pl.col("text").str.lengths() > 0)

for col_name, (start_pos, length) in schema_dict.items():
    scan = scan.with_columns(
        pl.col("text").str.slice(start_pos - 1, length).alias(col_name)
    )

# 最终收集结果
df = scan.drop("text").collect()

注意事项

  • 若你的schema中第二个参数是结束位置而非长度(比如(1,6)指第1到第6个字符),则截取长度应为 end_pos - start_pos + 1,切片逻辑调整为 str.slice(start_pos-1, end_pos - start_pos +1)
  • 可在提取列时通过.cast(pl.DataType)指定数据类型,例如.cast(pl.Float64)转换为浮点数

内容的提问来源于stack exchange,提问作者DBOak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 04:06:27