Polars Lazy DataFrame中用同行另一列填充列null值的可行性及性能咨询
Using Another Row's Column to Fill Nulls in Polars Lazy DataFrame
Absolutely feasible—and Polars makes this operation both intuitive and highly performant, even for large datasets. Here's how to pull it off, plus a breakdown of why the performance holds up:
Step-by-Step Implementation
The core tool here is Polars' fill_null method, which lets you reference another column directly to replace null values. Since you're working with a Lazy DataFrame, the operation will be optimized in Polars' query plan before execution, keeping things efficient.
Here's a concrete example:
import polars as pl # Create a sample Lazy DataFrame with nulls lazy_df = pl.LazyFrame({ "target_col": [2, None, 7, None, 11], "source_col": [20, 25, 30, 35, 40] }) # Fill nulls in target_col using values from source_col filled_lazy_df = lazy_df.with_columns( pl.col("target_col").fill_null(pl.col("source_col")) ) # Execute the lazy plan and get the result result = filled_lazy_df.collect() print(result)
This will output:
shape: (5, 2) ┌────────────┬────────────┐ │ target_col ┆ source_col │ │ --- ┆ --- │ │ i64 ┆ i64 │ ╞════════════╪════════════╡ │ 2 ┆ 20 │ │ 25 ┆ 25 │ │ 7 ┆ 30 │ │ 35 ┆ 35 │ │ 11 ┆ 40 │ └────────────┴────────────┘
Performance Breakdown
Polars is built for speed, and this operation leverages its key strengths:
- Vectorized Execution: Instead of slow row-by-row loops, Polars processes entire columns at once under the hood—this is the foundation of its speed with large datasets.
- Lazy Query Optimization: The fill operation gets merged into your overall query plan, so Polars avoids redundant computations. Using the Lazy API here doesn't add overhead; it just lets you defer execution until you're ready to collect results.
- Memory Efficiency: Polars uses contiguous memory blocks and minimizes unnecessary data copying, so even with massive datasets, this operation won't hog resources.
Quick Notes
- If your source column also has nulls, those will remain in the target column after filling. You can chain additional
fill_nullcalls (e.g.,fill_null(pl.col("source_col")).fill_null(0)) to handle edge cases like this. - You can overwrite the original column directly (no need for an alias) if that's your goal—just skip the
.alias(...)part in thewith_columnscall.
内容的提问来源于stack exchange,提问作者zareami10
相关产品推荐
相关产品推荐

