You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用np.where与for循环填充DataFrame时结果异常的问题求助

问题分析与解决方案

一、np.where方法的错误原因

你每次调用np.where时都直接覆盖了df['temp']的全部值,最后一行代码会把所有不满足19<=score<=40的行都设为'e',前面的条件判断完全被覆盖,导致最终所有值都是'e'。

正确做法是嵌套使用np.where,将后续判断放到前一个的else分支中,同时注意逻辑应该是「满足区间则赋值对应字母,否则进入下一层判断」:

import numpy as np

df['temp'] = np.where(df['nutrition-score-fr_100g'] <= -1, 'a',
                      np.where((df['nutrition-score-fr_100g'] >= 0) & (df['nutrition-score-fr_100g'] <= 2), 'b',
                               np.where((df['nutrition-score-fr_100g'] >= 3) & (df['nutrition-score-fr_100g'] <= 10), 'c',
                                        np.where((df['nutrition-score-fr_100g'] >= 11) & (df['nutrition-score-fr_100g'] <= 18), 'd',
                                                 np.where((df['nutrition-score-fr_100g'] >= 19) & (df['nutrition-score-fr_100g'] <= 40), 'e', None)))))

二、for循环方法的错误原因

  1. 循环中直接给df['temp']赋值会修改整个列,而非第i行,最终整个列的值会被最后一次循环满足的条件覆盖;
  2. range(0,3)是整数序列[0,1,2],如果你的nutrition-score-fr_100g是浮点数,in range永远不成立,会直接跳过该分支。

修正后的for循环写法:

# 先初始化temp列
df['temp'] = None
for i in range(len(df)):
    score = df['nutrition-score-fr_100g'].iloc[i]  # 用iloc访问单行更安全
    if score <= -1:
        df['temp'].iloc[i] = 'a'
    elif 0 <= score <= 2:
        df['temp'].iloc[i] = 'b'
    elif 3 <= score <= 10:
        df['temp'].iloc[i] = 'c'
    elif 11 <= score <= 18:
        df['temp'].iloc[i] = 'd'
    elif 19 <= score <= 40:
        df['temp'].iloc[i] = 'e'

更高效简洁的方案:使用pd.cut

针对这种区间划分场景,用pandas的pd.cut方法更高效,代码可读性更强:

import pandas as pd

# 定义区间边界和对应标签
bins = [-float('inf'), -1, 2, 10, 18, 40]
labels = ['a', 'b', 'c', 'd', 'e']
df['temp'] = pd.cut(df['nutrition-score-fr_100g'], bins=bins, labels=labels, include_lowest=True)
  • include_lowest=True确保左边界的数值(比如-1)被正确划分到对应组
  • 自动处理区间逻辑,无需手动嵌套判断

内容的提问来源于stack exchange,提问作者islem habibi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 01:31:04