You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Polars DataFrame转换为指定维度的NumPy数组并维持数据索引对应关系

如何将Polars DataFrame转换为指定维度的NumPy数组并维持数据索引对应关系

我目前遇到这样一个数据处理需求:手里有一个Polars DataFrame,包含300个流域(basin)的时序数据,每个流域恰好有10万条时间记录,每条记录对应40个变量,整体规模是3000万行、40个变量。我需要把它转换成形状为(300, 100000, 40)的NumPy数组,同时必须保证每个流域的时间序列、变量与原DataFrame的索引完全对应,不能出现错位的情况。

示例数据情况

举个具体的小例子,我的DataFrame结构和数据片段如下:

shape: (10, 7)
┌──────────────┬─────────────┬─────────────┬─────────────┬─────────────┬─────────────┬─────────────┐
│ HQprecipitat ┆ IRprecipita ┆ precipitati ┆ precipitati ┆ randomError ┆ basin_id    ┆ time        │
│ ion          ┆ tion        ┆ onCal       ┆ onUncal     ┆ ---         ┆ ---         ┆ ---         │
│ ---          ┆ ---         ┆ ---         ┆ ---         ┆ f32         ┆ str         ┆ datetime[μs │
│ f32          ┆ f32         ┆ f32         ┆ f32         ┆             ┆             ┆ ]           │
╞══════════════╪═════════════╪═════════════╪═════════════╪═════════════╪═════════════╪═════════════╡
│ null         ┆ null        ┆ null        ┆ null        ┆ null        ┆ anhui_62909 ┆ 1980-01-01  │
│              ┆             ┆             ┆             ┆             ┆ 400         ┆ 09:00:00    │
│ null         ┆ null        ┆ null        ┆ null        ┆ null        ┆ anhui_62909 ┆ 1980-01-01  │
│              ┆             ┆             ┆             ┆             ┆ 400         ┆ 12:00:00    │
│ null         ┆ null        ┆ null        ┆ null        ┆ null        ┆ anhui_62909 ┆ 1980-01-01  │
│              ┆             ┆             ┆             ┆             ┆ 400         ┆ 15:00:00    │
│ null         ┆ null        ┆ null        ┆ null        ┆ null        ┆ anhui_62909 ┆ 1980-01-01  │
│              ┆             ┆             ┆             ┆             ┆ 400         ┆ 18:00:00    │
│ null         ┆ null        ┆ null        ┆ null        ┆ null        ┆ anhui_62909 ┆ 1980-01-01  │
│              ┆             ┆             ┆             ┆             ┆ 400         ┆ 21:00:00    │
│ null         ┆ null        ┆ null        ┆ null        ┆ null        ┆ anhui_62909 ┆ 1980-01-02  │
│              ┆             ┆             ┆             ┆             ┆ 400         ┆ 00:00:00    │
│ null         ┆ null        ┆ null        ┆ null        ┆ null        ┆ anhui_62909 ┆ 1980-01-02  │
│              ┆             ┆             ┆             ┆             ┆ 400         ┆ 03:00:00    │
│ null         ┆ null        ┆ null        ┆ null        ┆ null        ┆ anhui_62909 ┆ 1980-01-02  │
│              ┆             ┆             ┆             ┆             ┆ 400         ┆ 06:00:00    │
│ null         ┆ null        ┆ null        ┆ null        ┆ null        ┆ anhui_62909 ┆ 1980-01-02  │
│              ┆             ┆             ┆             ┆             ┆ 400         ┆ 09:00:00    │
│ null         ┆ null        ┆ null        ┆ null        ┆ null        ┆ anhui_62909 ┆ 1980-01-02  │
│              ┆             ┆             ┆             ┆             ┆ 400         ┆ 12:00:00    │
└──────────────┴─────────────┴─────────────┴─────────────┴─────────────┴─────────────┴─────────────┘
# 目标是将其转换成形状为(1, 10, 5)的NumPy数组
# 其中1是流域数量,10是时间记录数,5是实际特征变量数(排除basin_id和time)

解决方案

这里提供两种可靠的方法,你可以根据自己的数据情况选择:

方法一:高效Reshape法(适用于数据规整的场景)

如果你的数据满足每个流域的时间记录数完全一致(都是10万条),且排序后同一流域的所有行是连续排列的,这种方法效率最高,直接通过reshape完成转换:

import polars as pl
import numpy as np

# 1. 先按流域和时间排序,确保时序正确且同流域数据连续
df_sorted = df.sort(["basin_id", "time"])

# 2. 筛选出需要纳入数组的特征列(排除basin_id和time)
feature_cols = [col for col in df_sorted.columns if col not in ["basin_id", "time"]]

# 3. 将特征列直接转为NumPy数组
features_np = df_sorted.select(feature_cols).to_numpy()

# 4. 按目标形状reshape
num_basins = 300
time_steps = 100000
num_features = len(feature_cols)
final_array = features_np.reshape(num_basins, time_steps, num_features)

# 验证形状是否符合要求
print(f"最终数组形状:{final_array.shape}")  # 应该输出(300, 100000, 40)

方法二:分组聚合法(更灵活,适用于数据可能有波动的场景)

如果担心数据可能存在小的异常(比如个别流域记录数有差异,或者同流域数据不连续),可以用分组聚合的方式,确保每个流域的数据单独处理后再组合:

import polars as pl
import numpy as np

# 1. 同样先排序保证时序
df_sorted = df.sort(["basin_id", "time"])

# 2. 筛选特征列
feature_cols = [col for col in df_sorted.columns if col not in ["basin_id", "time"]]

# 3. 按流域分组,将每个组的特征列转为二维数组
grouped_df = df_sorted.group_by("basin_id", maintain_order=True).agg(
    # 对每个特征列组,转成NumPy数组
    pl.col(feature_cols).apply(lambda group: group.to_numpy())
)

# 4. 将所有流域的数组合并成三维数组
final_array = np.stack(grouped_df["apply"].to_list())

# 验证形状
print(f"最终数组形状:{final_array.shape}")

关键注意事项

  • 排序是核心:无论用哪种方法,都必须先按basin_id和time排序,否则时序会混乱,数组里的索引对应关系就错了。
  • 缺失值处理:如果DataFrame里有null值(比如示例中的空值),转成NumPy数组后会变成np.nan(针对浮点类型)。如果需要填充缺失值,可以在排序后加上df_sorted = df_sorted.fill_null(0.0)(用0填充,或者根据业务逻辑选其他值)。
  • maintain_order=True的作用:在分组时加上这个参数,能保证最终数组里的流域顺序和原DataFrame中流域首次出现的顺序一致,避免流域索引错位。

备注:内容来源于stack exchange,提问作者forestbat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 17:52:56