You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中基于收入代理条件创建Income_Status列?

解决Pandas中基于收入代理值分类新列的问题

Got it, let's work through this together. First, I spotted a likely typo in your classification rules: the remaining cases (values that don't fall into the high or low thresholds) should probably be 中等平均收入 (medium average income, marked as 2) instead of low—since you already defined low for values above the upper threshold. I'll proceed with that correction, but feel free to adjust if that's not what you intended.

明确分类规则(修正后)

  • 1 = 高平均收入: Income_Proxy <= Average_Income_Proxy - 0.5 * Standard_dev_Income_Proxy
  • 3 = 低平均收入: Income_Proxy >= Average_Income_Proxy + 0.5 * Standard_dev_Income_Proxy
  • 2 = 中等平均收入: 所有不满足上述两个条件的情况

Pandas实现代码

We'll use numpy.select() here because it's efficient for multi-condition assignments, especially with larger datasets. Here's the full workflow:

1. 导入库并创建示例数据集

import pandas as pd
import numpy as np

# 你的原始数据集
data = {
    'Customer_id': ['123', '456', '789', '096', '158'],
    'Income_Proxy': [7681559.15, 15156.29, 50497.69, 44138.41, 67866.45],
    'Average_Income_Proxy': [44288.02]*5,
    'Standard_dev_Income_Proxy': [176568.76]*5
}
df = pd.DataFrame(data)

2. 计算分类阈值

# 计算上下阈值,直接用列运算实现向量化计算
lower_threshold = df['Average_Income_Proxy'] - 0.5 * df['Standard_dev_Income_Proxy']
upper_threshold = df['Average_Income_Proxy'] + 0.5 * df['Standard_dev_Income_Proxy']

3. 应用条件创建新列

# 定义条件列表和对应的标签
conditions = [
    df['Income_Proxy'] <= lower_threshold,
    df['Income_Proxy'] >= upper_threshold
]
labels = [1, 3]

# 创建Income_Status列,默认值为2(中等收入)
df['Income_Status'] = np.select(conditions, labels, default=2)

4. 查看结果

运行print(df)后会得到:

Customer_id  Income_Proxy  Average_Income_Proxy  Standard_dev_Income_Proxy  Income_Status
0         123    7681559.15               44288.02                  176568.76              3
1         456      15156.29               44288.02                  176568.76              1
2         789      50497.69               44288.02                  176568.76              2
3         096      44138.41               44288.02                  176568.76              2
4         158      67866.45               44288.02                  176568.76              2

如果你的原始规则确实是其余情况为低收入

If you meant that remaining values should also be marked as 3 (low average income), just change the default parameter to 3:

df['Income_Status'] = np.select(conditions, labels, default=3)

内容的提问来源于stack exchange,提问作者Saara Ligamena

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:02:04