如何在Polars中实现类似pandas.reindex(method='ffill')的时间列功能?
在Polars中实现保留原空值的索引扩容与向前填充
问题背景
在Pandas中,可通过reindex结合method="ffill"实现:添加新索引行时对新增行向前填充数值,但保留原DataFrame中的空值,示例代码与输出如下:
Pandas 实现示例
import numpy as np import pandas as pd df = pd.DataFrame(data={"a": [1.0, 2.0, np.nan, 3.0]}, index=pd.date_range("2020", periods=4, freq="T")) print("原DataFrame:") print(df) df = df.reindex(index=df.index.union(pd.date_range("2020-01-01 00:01:30", periods=2, freq="T")), method="ffill") print("\n扩容并填充后的DataFrame:") print(df)
输出结果:
原DataFrame: a 2020-01-01 00:00:00 1.0 2020-01-01 00:01:00 2.0 2020-01-01 00:02:00 NaN 2020-01-01 00:03:00 3.0 扩容并填充后的DataFrame: a 2020-01-01 00:00:00 1.0 2020-01-01 00:01:00 2.0 2020-01-01 00:01:30 2.0 2020-01-01 00:02:00 NaN 2020-01-01 00:02:30 NaN 2020-01-01 00:03:00 3.0
需求是在Polars中实现相同效果,且必须保证性能(选择Polars正是看重其性能优势)。此前尝试的concat->sort->ffill->unique方案会填充原数据中的空值,不符合要求。
Polars 实现方案
核心思路:通过标记原数据行,仅对新增行执行向前填充,全程使用惰性API保证性能。
Python 版本实现
import polars as pl from datetime import datetime # 构建原DataFrame df = pl.DataFrame( data={"a": [1.0, 2.0, None, 3.0]}, index=pl.date_range(datetime(2020,1,1), periods=4, interval="1m") ).rename({"index": "timestamp"}) # 定义新增的索引范围 new_timestamps = pl.date_range(datetime(2020,1,1,0,1,30), periods=2, interval="1m") # 合并原数据与新增索引,标记原数据行 combined = ( pl.concat( [ df.with_columns(pl.lit(True).alias("is_original")), pl.DataFrame({"timestamp": new_timestamps, "a": None, "is_original": False}) ] ) .sort("timestamp") ) # 仅对非原数据行进行向前填充,保留原数据的空值 result = ( combined .with_columns( pl.when(pl.col("is_original")) .then(pl.col("a")) .otherwise(pl.col("a").forward_fill()) .alias("a") ) .drop("is_original") ) print("原DataFrame:") print(df) print("\n扩容并填充后的DataFrame:") print(result)
输出结果:
原DataFrame: ┌─────────────────────┬──────┐ │ timestamp ┆ a │ │ --- ┆ --- │ │ datetime[μs] ┆ f64 │ ╞═════════════════════╪══════╡ │ 2020-01-01 00:00:00 ┆ 1.0 │ │ 2020-01-01 00:01:00 ┆ 2.0 │ │ 2020-01-01 00:02:00 ┆ null │ │ 2020-01-01 00:03:00 ┆ 3.0 │ └─────────────────────┴──────┘ 扩容并填充后的DataFrame: ┌─────────────────────┬──────┐ │ timestamp ┆ a │ │ --- ┆ --- │ │ datetime[μs] ┆ f64 │ ╞═════════════════════╪══════╡ │ 2020-01-01 00:00:00 ┆ 1.0 │ │ 2020-01-01 00:01:00 ┆ 2.0 │ │ 2020-01-01 00:01:30 ┆ 2.0 │ │ 2020-01-01 00:02:00 ┆ null │ │ 2020-01-01 00:02:30 ┆ null │ │ 2020-01-01 00:03:00 ┆ 3.0 │ └─────────────────────┴──────┘
Rust 版本实现
针对你提供的Rust代码修改,实现仅填充新增行、保留原空值的逻辑:
use polars::prelude::*; use chrono::{DateTime, Utc}; fn main() -> PolarsResult<()> { // 构建原DataFrame let timestamps = vec![ DateTime::parse_from_rfc3339("2020-01-01T00:00:00Z")?.with_timezone(&Utc), DateTime::parse_from_rfc3339("2020-01-01T00:01:00Z")?.with_timezone(&Utc), DateTime::parse_from_rfc3339("2020-01-01T00:02:00Z")?.with_timezone(&Utc), DateTime::parse_from_rfc3339("2020-01-01T00:03:00Z")?.with_timezone(&Utc), ]; let a = Series::new("a", vec![Some(1.0), Some(2.0), None, Some(3.0)]); let mut df = DataFrame::new(vec![Series::new("timestamp", timestamps), a])?; // 定义新增索引 let new_timestamps = vec![ DateTime::parse_from_rfc3339("2020-01-01T00:01:30Z")?.with_timezone(&Utc), DateTime::parse_from_rfc3339("2020-01-01T00:02:30Z")?.with_timezone(&Utc), ]; let new_rows = DataFrame::new(vec![ Series::new("timestamp", new_timestamps), Series::full_null("a", 2, &DataType::Float64)?, ])?; // 合并并标记原数据行 let combined = df .lazy() .with_column(lit(true).alias("is_original")) .concat(&[new_rows.lazy().with_column(lit(false).alias("is_original"))]) .sort(["timestamp"], false) .collect()?; // 仅对非原数据行执行向前填充 let result = combined .lazy() .with_column( when(col("is_original")) .then(col("a")) .otherwise(col("a").forward_fill(None)) .alias("a"), ) .drop("is_original") .collect()?; println!("原DataFrame:\n{}", df); println!("\n扩容并填充后的DataFrame:\n{}", result); Ok(()) }
方案说明
- 通过
is_original标记区分原数据行与新增行,确保仅对新增行应用forward_fill - 全程使用Polars惰性API(
lazy()),避免不必要的内存开销,保证高性能 - 最终结果完全匹配Pandas的行为:新增行向前填充,原数据的空值完整保留
内容的提问来源于stack exchange,提问作者Are Haartveit
相关产品推荐
相关产品推荐

