如何使用Pandas replace函数仅替换指定列的yes/no值?
解决Pandas仅替换指定列yes/no值的问题
问题场景
读取数据集:
dataset = pd.read_csv('./file.csv') dataset.head()
得到如下数据:
age sex smoker married region price 0 39 female yes no us 250000 1 28 male no no us 400000 2 23 male no yes europe 389000 3 17 male no no asia 230000 4 43 male no yes asia 243800
需要将smoker列的yes/no替换为1/0,但不修改married列的yes/no值,且优先使用Pandas的replace函数。此前对整个DataFrame调用replace会修改所有列的对应值:
dataset = dataset.replace(to_replace='yes', value='1') dataset = dataset.replace(to_replace='no', value='0')
执行后married列的yes/no也被修改,不符合需求。
解决方案
直接对目标列smoker单独调用replace方法,传入映射字典即可:
# 仅替换smoker列的yes/no为1/0 dataset['smoker'] = dataset['smoker'].replace({'yes': 1, 'no': 0})
执行后得到的结果:
age sex smoker married region price 0 39 female 1 no us 250000 1 28 male 0 no us 400000 2 23 male 0 yes europe 389000 3 17 male 0 no asia 230000 4 43 male 0 yes asia 243800
补充写法
也可以通过loc定位列实现,效果一致:
dataset.loc[:, 'smoker'] = dataset.loc[:, 'smoker'].replace({'yes': 1, 'no': 0})
核心逻辑是只对目标列调用replace方法,而非对整个DataFrame操作,这样就不会影响其他列的内容。
内容的提问来源于stack exchange,提问作者wiwa1978
相关产品推荐
相关产品推荐

