使用np.where与for循环填充DataFrame时结果异常的问题求助
问题分析与解决方案
一、np.where方法的错误原因
你每次调用np.where时都直接覆盖了df['temp']的全部值,最后一行代码会把所有不满足19<=score<=40的行都设为'e',前面的条件判断完全被覆盖,导致最终所有值都是'e'。
正确做法是嵌套使用np.where,将后续判断放到前一个的else分支中,同时注意逻辑应该是「满足区间则赋值对应字母,否则进入下一层判断」:
import numpy as np df['temp'] = np.where(df['nutrition-score-fr_100g'] <= -1, 'a', np.where((df['nutrition-score-fr_100g'] >= 0) & (df['nutrition-score-fr_100g'] <= 2), 'b', np.where((df['nutrition-score-fr_100g'] >= 3) & (df['nutrition-score-fr_100g'] <= 10), 'c', np.where((df['nutrition-score-fr_100g'] >= 11) & (df['nutrition-score-fr_100g'] <= 18), 'd', np.where((df['nutrition-score-fr_100g'] >= 19) & (df['nutrition-score-fr_100g'] <= 40), 'e', None)))))
二、for循环方法的错误原因
- 循环中直接给
df['temp']赋值会修改整个列,而非第i行,最终整个列的值会被最后一次循环满足的条件覆盖; range(0,3)是整数序列[0,1,2],如果你的nutrition-score-fr_100g是浮点数,in range永远不成立,会直接跳过该分支。
修正后的for循环写法:
# 先初始化temp列 df['temp'] = None for i in range(len(df)): score = df['nutrition-score-fr_100g'].iloc[i] # 用iloc访问单行更安全 if score <= -1: df['temp'].iloc[i] = 'a' elif 0 <= score <= 2: df['temp'].iloc[i] = 'b' elif 3 <= score <= 10: df['temp'].iloc[i] = 'c' elif 11 <= score <= 18: df['temp'].iloc[i] = 'd' elif 19 <= score <= 40: df['temp'].iloc[i] = 'e'
更高效简洁的方案:使用pd.cut
针对这种区间划分场景,用pandas的pd.cut方法更高效,代码可读性更强:
import pandas as pd # 定义区间边界和对应标签 bins = [-float('inf'), -1, 2, 10, 18, 40] labels = ['a', 'b', 'c', 'd', 'e'] df['temp'] = pd.cut(df['nutrition-score-fr_100g'], bins=bins, labels=labels, include_lowest=True)
include_lowest=True确保左边界的数值(比如-1)被正确划分到对应组- 自动处理区间逻辑,无需手动嵌套判断
内容的提问来源于stack exchange,提问作者islem habibi
相关产品推荐
相关产品推荐

