Polars DataFrame上采样时如何仅对部分列执行向前填充?
Polars上采样时如何仅对部分列执行forward_fill?
在使用Polars对DataFrame进行上采样时,能否仅对部分列执行向前填充(forward_fill)?
举例来说:对示例DataFrame补全缺失日期,当前通过upsample结合pl.all().forward_fill()可以实现全列向前填充,处理大数据集速度很快,输出符合预期。但现在希望排除utc_time列,让该列在填充后保留空值而非被填充。尝试用pl.col(...)替换pl.all()指定列时,未指定的列会被移除,请问该怎么实现?
示例代码
import polars as pl from datetime import datetime df = pl.DataFrame( { 'utc_created':pl.date_range(datetime(2021, 12, 16), datetime(2021, 12, 22, 0), interval="2d", eager=True), 'utc_time':['21:12:06','21:20:06','17:51:10','03:54:49'], 'sku':[9100000801,9100000801,9100000801,9100000801], 'old':[18,17,16,15], 'new':[17,16,15,14], 'alert_type':['Inventory','Inventory','Inventory','Inventory'], 'alert_level':['Info','Info','Info','Info'] } ) # 当前全列向前填充的实现代码 df = (df.upsample(time_column='utc_created',every='1d', group_by='sku') .select(pl.all().forward_fill()) )
当前输出
shape: (7, 7) ┌─────────────┬──────────┬────────────┬─────┬─────┬────────────┬─────────────┐ │ utc_created ┆ utc_time ┆ sku ┆ old ┆ new ┆ alert_type ┆ alert_level │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ date ┆ str ┆ i64 ┆ i64 ┆ i64 ┆ str ┆ str │ ╞═════════════╪══════════╪════════════╪═════╪═════╪════════════╪═════════════╡ │ 2021-12-16 ┆ 21:12:06 ┆ 9100000801 ┆ 18 ┆ 17 ┆ Inventory ┆ Info │ │ 2021-12-17 ┆ 21:12:06 ┆ 9100000801 ┆ 18 ┆ 17 ┆ Inventory ┆ Info │ │ 2021-12-18 ┆ 21:20:06 ┆ 9100000801 ┆ 17 ┆ 16 ┆ Inventory ┆ Info │ │ 2021-12-19 ┆ 21:20:06 ┆ 9100000801 ┆ 17 ┆ 16 ┆ Inventory ┆ Info │ │ 2021-12-20 ┆ 17:51:10 ┆ 9100000801 ┆ 16 ┆ 15 ┆ Inventory ┆ Info │ │ 2021-12-21 ┆ 17:51:10 ┆ 9100000801 ┆ 16 ┆ 15 ┆ Inventory ┆ Info │ │ 2021-12-22 ┆ 03:54:49 ┆ 9100000801 ┆ 15 ┆ 14 ┆ Inventory ┆ Info │ └─────────────┴──────────┴────────────┴─────┴─────┴────────────┴─────────────┘
解决方案
要实现仅对指定列执行forward_fill,同时保留未指定的列(比如utc_time),可以通过组合列选择规则实现:
- 对需要填充的列应用
forward_fill() - 直接保留不需要填充的列(不做任何填充处理)
优化后代码
df = (df.upsample(time_column='utc_created', every='1d', group_by='sku') .select( # 对除utc_time外的所有列执行向前填充 pl.all().exclude('utc_time').forward_fill(), # 直接保留utc_time列,不做填充 pl.col('utc_time') ) )
优化后输出
shape: (7, 7) ┌─────────────┬────────────┬─────┬─────┬────────────┬─────────────┬──────────┐ │ utc_created ┆ sku ┆ old ┆ new ┆ alert_type ┆ alert_level ┆ utc_time │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ date ┆ i64 ┆ i64 ┆ i64 ┆ str ┆ str ┆ str │ ╞═════════════╪════════════╪═════╪═════╪════════════╪═════════════╪══════════╡ │ 2021-12-16 ┆ 9100000801 ┆ 18 ┆ 17 ┆ Inventory ┆ Info ┆ 21:12:06 │ │ 2021-12-17 ┆ 9100000801 ┆ 18 ┆ 17 ┆ Inventory ┆ Info ┆ null │ │ 2021-12-18 ┆ 9100000801 ┆ 17 ┆ 16 ┆ Inventory ┆ Info ┆ 21:20:06 │ │ 2021-12-19 ┆ 9100000801 ┆ 17 ┆ 16 ┆ Inventory ┆ Info ┆ null │ │ 2021-12-20 ┆ 9100000801 ┆ 16 ┆ 15 ┆ Inventory ┆ Info ┆ 17:51:10 │ │ 2021-12-21 ┆ 9100000801 ┆ 16 ┆ 15 ┆ Inventory ┆ Info ┆ null │ │ 2021-12-22 ┆ 9100000801 ┆ 15 ┆ 14 ┆ Inventory ┆ Info ┆ 03:54:49 │ └─────────────┴────────────┴─────┴─────┴────────────┴─────────────┴──────────┘
补充说明
- 如果需要指定具体填充列而非排除列,可直接列出目标列:
df = (df.upsample(time_column='utc_created', every='1d', group_by='sku') .select( pl.col(['sku', 'old', 'new', 'alert_type', 'alert_level', 'utc_created']).forward_fill(), pl.col('utc_time') ) ) - 该方式不会移除任何列,同时能精准控制填充范围,完全适配大数据集的高效处理需求。
内容的提问来源于stack exchange,提问作者DBOak
相关产品推荐
相关产品推荐

