You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python Polars中比较行日期值并处理格式差异?

解决Polars DataFrame中出生日期月份重复标记问题

核心思路

问题根源在于字符串格式的日期中,月份的前导零导致被误判为不同值。解决关键是先将字符串日期解析为标准日期类型,统一提取月份的数字表示,再根据月份是否重复进行标记。

具体代码实现

import polars as pl

# 初始化原始DataFrame
df = pl.DataFrame({
    'idx': [1,2,3,4,5,6],
    'date_of_birth': ['03/06/1990','3/06/1990','11/12/2000','01/02/2021','1/02/2021','3/06/1990']
})

# 处理流程:解析日期→提取月份→标记重复月份
result_df = df.with_columns(
    # 解析日期,自动处理前导零差异
    parsed_date=pl.col('date_of_birth').str.to_date(format="%m/%d/%Y"),
    # 提取整数格式的月份,统一表示
    month=pl.col('parsed_date').dt.month()
).with_columns(
    # 标记重复月份:出现次数≥2则标记为yes,否则为no
    is_same_month=pl.when(pl.col('month').is_duplicated()).then('yes').otherwise('no')
).drop('parsed_date')  # 可选:删除中间解析列

print(result_df)

代码说明

  • 日期解析:str.to_date(format="%m/%d/%Y")会自动识别带前导零和不带前导零的月份(如"03"和"3"都会被解析为3月),将字符串转为标准日期类型。
  • 提取月份:dt.month()从解析后的日期中提取整数月份,彻底消除前导零带来的格式差异。
  • 标记重复:is_duplicated()判断当前行的月份是否在DataFrame中重复出现,重复则标记为yes,否则为no。

输出结果

执行代码后得到如下结果:

shape: (6, 4)
┌─────┬────────────────┬───────┬──────────────┐
│ idx ┆ date_of_birth  ┆ month ┆ is_same_month│
│ --- ┆ ---            ┆ ---   ┆ ---          │
│ i64 ┆ str            ┆ i32   ┆ str          │
╞═════╪════════════════╪═══════╪══════════════╡
│ 1   ┆ 03/06/1990     ┆ 3     ┆ yes          │
│ 2   ┆ 3/06/1990      ┆ 3     ┆ yes          │
│ 3   ┆ 11/12/2000     ┆ 11    ┆ no           │
│ 4   ┆ 01/02/2021     ┆ 1     ┆ yes          │
│ 5   ┆ 1/02/2021      ┆ 1     ┆ yes          │
│ 6   ┆ 3/06/1990      ┆ 3     ┆ yes          │
└─────┴────────────────┴───────┴──────────────┘

内容的提问来源于stack exchange,提问作者myamulla_ciencia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 19:27:36