基于动态时间索引条件为Pandas DataFrame新增列的实现方案
动态批量处理多时段条件为Pandas DataFrame新增列
需求说明
基于时间索引为Pandas DataFrame(df)新增一列:
- 当索引处于
years_list中任意年份的3月1日至3月7日期间时,新列值为1 - 其余情况值为2
years_list支持动态调整(可包含任意数量年份)
现有代码问题
原代码存在逻辑错误,导致新列与df长度不匹配,且仅能匹配最后一个年份的日期范围:
new_col= [] years_list= [2020,2021,2023] for year in years_list: start_date= pd.to_datetime(str(year)+'-03-01 00:00:00') end_date = pd.to_datetime(str(year)+'-03-07 00:00:00') for idx in range(len(df)): if df.index[idx] >= start_date and df.index[idx] <= end_date: new_col.append(1) else: new_col.append(2) df["newC"] = new_col
错误原因:
- 第一个循环结束后,
start_date和end_date仅保留最后一个年份(2023)的日期值 - 遍历索引的逻辑无法同时匹配多个年份的日期范围
解决方案
方法1:向量化操作(推荐,高效简洁)
直接通过索引的年份和月日属性构建条件,无需循环,完美支持动态年份列表:
import pandas as pd years_list = [2020, 2021, 2023] # 提取索引的年份和月日字符串(格式如'03-01') idx_year = df.index.year idx_md = df.index.strftime('%m-%d') # 构建复合条件:年份在目标列表中,且日期处于3月1日至7日之间 condition = (idx_year.isin(years_list)) & (idx_md >= '03-01') & (idx_md <= '03-07') # 为新列赋值 df['newC'] = condition.map({True: 1, False: 2})
方法2:动态生成目标时间范围集合
如果需要精确匹配包含时分秒的时间戳,可先生成所有目标时段的时间集合,再判断索引是否属于该集合:
import pandas as pd years_list = [2020, 2021, 2023] target_ranges = [] # 生成所有目标年份的3月1日至3月7日完整时间范围 for year in years_list: start = pd.to_datetime(f'{year}-03-01 00:00:00') end = pd.to_datetime(f'{year}-03-07 23:59:59') # 覆盖3月7日全天 target_ranges.append(pd.date_range(start, end, freq='S')) # 按秒生成,可按需调整频率 # 合并所有目标时间戳为一个集合 target_timestamps = pd.concat(target_ranges) # 判断索引是否在目标集合中并赋值 df['newC'] = df.index.isin(target_timestamps).map({True: 1, False: 2})
预期结果示例
| date time (index) | new col |
|---|---|
| 2020-01-01 00:00:00 | 2 |
| ... | 2 |
| 2020-02-29 00:00:00 | 2 |
| 2020-03-01 00:00:00 | 1 |
| 2020-03-02 00:00:00 | 1 |
| ... | 1 |
| 2020-03-07 00:00:00 | 1 |
| 2020-03-08 00:00:00 | 2 |
| ... | ... |
| 2023-03-07 00:00:00 | 1 |
| 2023-03-08 00:00:00 | 2 |
内容的提问来源于stack exchange,提问作者jess
相关产品推荐
相关产品推荐

