如何将Polars DataFrame转换为指定维度的NumPy数组并维持数据索引对应关系
如何将Polars DataFrame转换为指定维度的NumPy数组并维持数据索引对应关系
我目前遇到这样一个数据处理需求:手里有一个Polars DataFrame,包含300个流域(basin)的时序数据,每个流域恰好有10万条时间记录,每条记录对应40个变量,整体规模是3000万行、40个变量。我需要把它转换成形状为(300, 100000, 40)的NumPy数组,同时必须保证每个流域的时间序列、变量与原DataFrame的索引完全对应,不能出现错位的情况。
示例数据情况
举个具体的小例子,我的DataFrame结构和数据片段如下:
shape: (10, 7) ┌──────────────┬─────────────┬─────────────┬─────────────┬─────────────┬─────────────┬─────────────┐ │ HQprecipitat ┆ IRprecipita ┆ precipitati ┆ precipitati ┆ randomError ┆ basin_id ┆ time │ │ ion ┆ tion ┆ onCal ┆ onUncal ┆ --- ┆ --- ┆ --- │ │ --- ┆ --- ┆ --- ┆ --- ┆ f32 ┆ str ┆ datetime[μs │ │ f32 ┆ f32 ┆ f32 ┆ f32 ┆ ┆ ┆ ] │ ╞══════════════╪═════════════╪═════════════╪═════════════╪═════════════╪═════════════╪═════════════╡ │ null ┆ null ┆ null ┆ null ┆ null ┆ anhui_62909 ┆ 1980-01-01 │ │ ┆ ┆ ┆ ┆ ┆ 400 ┆ 09:00:00 │ │ null ┆ null ┆ null ┆ null ┆ null ┆ anhui_62909 ┆ 1980-01-01 │ │ ┆ ┆ ┆ ┆ ┆ 400 ┆ 12:00:00 │ │ null ┆ null ┆ null ┆ null ┆ null ┆ anhui_62909 ┆ 1980-01-01 │ │ ┆ ┆ ┆ ┆ ┆ 400 ┆ 15:00:00 │ │ null ┆ null ┆ null ┆ null ┆ null ┆ anhui_62909 ┆ 1980-01-01 │ │ ┆ ┆ ┆ ┆ ┆ 400 ┆ 18:00:00 │ │ null ┆ null ┆ null ┆ null ┆ null ┆ anhui_62909 ┆ 1980-01-01 │ │ ┆ ┆ ┆ ┆ ┆ 400 ┆ 21:00:00 │ │ null ┆ null ┆ null ┆ null ┆ null ┆ anhui_62909 ┆ 1980-01-02 │ │ ┆ ┆ ┆ ┆ ┆ 400 ┆ 00:00:00 │ │ null ┆ null ┆ null ┆ null ┆ null ┆ anhui_62909 ┆ 1980-01-02 │ │ ┆ ┆ ┆ ┆ ┆ 400 ┆ 03:00:00 │ │ null ┆ null ┆ null ┆ null ┆ null ┆ anhui_62909 ┆ 1980-01-02 │ │ ┆ ┆ ┆ ┆ ┆ 400 ┆ 06:00:00 │ │ null ┆ null ┆ null ┆ null ┆ null ┆ anhui_62909 ┆ 1980-01-02 │ │ ┆ ┆ ┆ ┆ ┆ 400 ┆ 09:00:00 │ │ null ┆ null ┆ null ┆ null ┆ null ┆ anhui_62909 ┆ 1980-01-02 │ │ ┆ ┆ ┆ ┆ ┆ 400 ┆ 12:00:00 │ └──────────────┴─────────────┴─────────────┴─────────────┴─────────────┴─────────────┴─────────────┘ # 目标是将其转换成形状为(1, 10, 5)的NumPy数组 # 其中1是流域数量,10是时间记录数,5是实际特征变量数(排除basin_id和time)
解决方案
这里提供两种可靠的方法,你可以根据自己的数据情况选择:
方法一:高效Reshape法(适用于数据规整的场景)
如果你的数据满足每个流域的时间记录数完全一致(都是10万条),且排序后同一流域的所有行是连续排列的,这种方法效率最高,直接通过reshape完成转换:
import polars as pl import numpy as np # 1. 先按流域和时间排序,确保时序正确且同流域数据连续 df_sorted = df.sort(["basin_id", "time"]) # 2. 筛选出需要纳入数组的特征列(排除basin_id和time) feature_cols = [col for col in df_sorted.columns if col not in ["basin_id", "time"]] # 3. 将特征列直接转为NumPy数组 features_np = df_sorted.select(feature_cols).to_numpy() # 4. 按目标形状reshape num_basins = 300 time_steps = 100000 num_features = len(feature_cols) final_array = features_np.reshape(num_basins, time_steps, num_features) # 验证形状是否符合要求 print(f"最终数组形状:{final_array.shape}") # 应该输出(300, 100000, 40)
方法二:分组聚合法(更灵活,适用于数据可能有波动的场景)
如果担心数据可能存在小的异常(比如个别流域记录数有差异,或者同流域数据不连续),可以用分组聚合的方式,确保每个流域的数据单独处理后再组合:
import polars as pl import numpy as np # 1. 同样先排序保证时序 df_sorted = df.sort(["basin_id", "time"]) # 2. 筛选特征列 feature_cols = [col for col in df_sorted.columns if col not in ["basin_id", "time"]] # 3. 按流域分组,将每个组的特征列转为二维数组 grouped_df = df_sorted.group_by("basin_id", maintain_order=True).agg( # 对每个特征列组,转成NumPy数组 pl.col(feature_cols).apply(lambda group: group.to_numpy()) ) # 4. 将所有流域的数组合并成三维数组 final_array = np.stack(grouped_df["apply"].to_list()) # 验证形状 print(f"最终数组形状:{final_array.shape}")
关键注意事项
- 排序是核心:无论用哪种方法,都必须先按
basin_id和time排序,否则时序会混乱,数组里的索引对应关系就错了。 - 缺失值处理:如果DataFrame里有
null值(比如示例中的空值),转成NumPy数组后会变成np.nan(针对浮点类型)。如果需要填充缺失值,可以在排序后加上df_sorted = df_sorted.fill_null(0.0)(用0填充,或者根据业务逻辑选其他值)。 maintain_order=True的作用:在分组时加上这个参数,能保证最终数组里的流域顺序和原DataFrame中流域首次出现的顺序一致,避免流域索引错位。
备注:内容来源于stack exchange,提问作者forestbat
相关产品推荐
相关产品推荐

