You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决不同长度列表逐元素比较问题:为DataFrame添加日期匹配标记列

为DataFrame新增日期匹配标记列

原始数据

DataFrame:

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'datetime': ['2023-01-01 12:00:00', '2023-01-02 12:00:00', '2023-01-03 12:00:00',
                 '2023-01-04 12:00:00', '2023-01-05 12:00:00', '2023-01-06 12:00:00',
                 '2023-01-07 12:00:00'],
    'col1': [100, 120, 140, 160, 200, 430, 890],
    'col2': [200, 400, 500, 700, 300, 200, 100]
})
# 先将datetime列转为datetime类型
df['datetime'] = pd.to_datetime(df['datetime'])

日期列表:

dates = ["2023-01-01", "2023-01-03", "2023-01-07"]

需求

新增一列col3,当df['datetime']的日期部分与dates列表中的元素匹配时填充1,否则填充0。

问题分析

你之前的代码报错是因为两个核心问题:

  1. np.isin(dates, pd.DatetimeIndex(df['datetime']).date)参数顺序错误,应该判断df的日期是否在dates列表中,而非反过来,否则得到的结果长度和df不匹配;
  2. 执行时col3尚未创建,不能直接使用df['col3']==1这类写法。

解决方案

方法1:isin+类型转换

# 将dates转为date对象集合
target_dates = pd.to_datetime(dates).date
# 提取df中datetime的日期部分,判断是否在目标日期里,转成int类型(True→1,False→0)
df['col3'] = df['datetime'].dt.date.isin(target_dates).astype(int)

方法2:使用np.where

target_dates = pd.to_datetime(dates).date
df['col3'] = np.where(df['datetime'].dt.date.isin(target_dates), 1, 0)

最终结果

执行代码后打印df,输出如下:

datetime  col1  col2  col3
0 2023-01-01 12:00:00   100   200     1
1 2023-01-02 12:00:00   120   400     0
2 2023-01-03 12:00:00   140   500     1
3 2023-01-04 12:00:00   160   700     0
4 2023-01-05 12:00:00   200   300     0
5 2023-01-06 12:00:00   430   200     0
6 2023-01-07 12:00:00   890   100     1

内容的提问来源于stack exchange,提问作者Pythoneer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 21:00:51