You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中基于多列表if/elif条件创建土壤质地分类新列?

问题分析与解决

原代码的问题

  1. 列表语法错误:coarse_tex = ['Sand', 'Loamy Sand' 'Sandy Loam']中,'Loamy Sand'与'Sandy Loam'之间缺少逗号,会被合并成一个错误字符串,导致无法正确匹配土壤质地。
  2. 函数逻辑错误:函数内用sha_df['texture'].isin(coarse_tex)是对整个列做判断,返回布尔值数组,而非针对传入的单个texture值判断,这会触发ValueError。
  3. apply使用不当:直接调用sha_df.apply(texture_classifier)未指定处理维度,无法正常逐行执行函数。

修正后的代码

第一步:修正分类列表

# 定义土壤质地分类列表
coarse_tex = ['Sand', 'Loamy Sand', 'Sandy Loam']
medium_tex = ['Loam', 'Silt Loam', 'Silt', 'Sandy Clay Loam']
fine_tex = ['Clay Loam', 'Silty Clay Loam', 'Sandy Clay', 'Silty Clay', 'Clay']

第二步:改进分类函数

函数需接收单个texture值,针对单个元素判断:

# 定义质地分类函数
def texture_classifier(texture):
    if texture in coarse_tex:
        return 'coarse'
    elif texture in medium_tex:
        return 'medium'
    elif texture in fine_tex:
        return 'fine'
    else:
        return None  # 处理不在分类列表中的异常值

第三步:应用函数生成新列

推荐两种实现方式:

# 方法1:单列apply(简洁直观)
sha_df['textural_class'] = sha_df['texture'].apply(texture_classifier)

# 方法2:使用lambda映射(逻辑等价)
sha_df['textural_class'] = sha_df['texture'].map(lambda x: texture_classifier(x))

大数据集高效方案:numpy.select

如果数据集规模较大,用np.select比apply速度更快:

import numpy as np

# 定义匹配条件与对应分类
conditions = [
    sha_df['texture'].isin(coarse_tex),
    sha_df['texture'].isin(medium_tex),
    sha_df['texture'].isin(fine_tex)
]
choices = ['coarse', 'medium', 'fine']

sha_df['textural_class'] = np.select(conditions, choices, default=None)

最终结果

处理后的数据集中会生成textural_class列,示例如下:

textural_class
coarse
medium
fine

内容的提问来源于stack exchange,提问作者Kwabena Addae Sarpong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 15:05:30