如何基于DataFrame某一列的数值规则修改另一列的对应字段值
实现方法
这里提供3种常用的实现方案,可根据你的数据规模和使用习惯选择:
方案1:np.select 实现(通用性最强,条件复杂时优先选)
import pandas as pd import numpy as np # 构造示例数据,实际使用时替换为你自己的DataFrame读入逻辑即可 df = pd.DataFrame({ 'val': [10,70,61,45,32,77,11], 'type': ['new','new','new','old','new','mid','mid'] }) # 定义匹配条件和对应返回值 conditions = [ (df['type'] == 'new') & (df['val'] <= 20), (df['type'] == 'new') & (df['val'] > 20) & (df['val'] < 50), (df['type'] == 'new') & (df['val'] >= 50) ] choices = ['new-1', 'new-2', 'new-3'] # 条件匹配不到的默认保留原type值 df['type'] = np.select(conditions, choices, default=df['type'])
方案2:pd.cut 分箱实现(数值区间类场景写法更简洁)
# 仅筛选type为new的行做处理,其余行不改动 mask = df['type'] == 'new' df.loc[mask, 'type'] = 'new-' + pd.cut( df.loc[mask, 'val'], bins=[-float('inf'), 20, 50, float('inf')], labels=['1', '2', '3'] ).astype(str)
方案3:apply 实现(仅适合小数据量使用,逻辑最直观)
def type_mapper(row): if row['type'] != 'new': return row['type'] if row['val'] <= 20: return 'new-1' elif 20 < row['val'] < 50: return 'new-2' return 'new-3' df['type'] = df.apply(type_mapper, axis=1)
运行以上任意一种方案后得到的结果都和你的预期一致。
内容的提问来源于stack exchange,提问作者skulldoger
相关产品推荐
相关产品推荐

