You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Stata中利用分位数创建收入分组变量?

按收入分位数创建分组变量的实现方法

R语言

先计算收入的50th和90th分位数,再通过条件判断生成分组:

# 计算分位数(忽略缺失值)
p50 <- quantile(df$income, 0.5, na.rm = TRUE)
p90 <- quantile(df$income, 0.9, na.rm = TRUE)

# 用dplyr的case_when创建分组
library(dplyr)
df$income_group <- case_when(
  df$income < p50 ~ 1,
  df$income >= p50 & df$income <= p90 ~ 2,
  df$income > p90 ~ 3,
  TRUE ~ NA_real_  # 处理缺失值
)

如果不用dplyr,也可以用基础R的嵌套ifelse:

df$income_group <- ifelse(df$income < p50, 1,
                          ifelse(df$income <= p90, 2, 3))

Python(Pandas)

方法1:直接用qcut按分位数分组并映射标签:

import pandas as pd

df['income_group'] = pd.qcut(
    df['income'],
    q=[0, 0.5, 0.9, 1],
    labels=[1, 2, 3]
)

方法2:先计算分位数,再用cut自定义区间:

p50 = df['income'].quantile(0.5)
p90 = df['income'].quantile(0.9)

df['income_group'] = pd.cut(
    df['income'],
    bins=[-float('inf'), p50, p90, float('inf')],
    labels=[1, 2, 3],
    include_lowest=True
)

Stata

方式1:用xtile生成等分组后重新编码:

// 先把收入分成10个等分组
xtile temp_group = income, nq(10)

// 重新编码为目标分组
recode temp_group (1/5=1) (6/9=2) (10=3), gen(income_group)

// 删除临时变量
drop temp_group

方式2:直接计算分位数后判断赋值:

egen p50 = pctile(income), p(50)
egen p90 = pctile(income), p(90)

gen income_group = 1 if income < p50
replace income_group = 2 if income >= p50 & income <= p90
replace income_group = 3 if income > p90

内容的提问来源于stack exchange,提问作者Toxinator

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 21:32:42