You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用data.table对长格式数据按分组行计算失业率

用data.table实现分组失业率计算(行形式输出)

针对你的长格式数据需求,这里提供两种data.table的实现方式,都能完成按sex、region、data_type分组计算失业率,并输出行形式结果:

方法1:宽格式转换计算后转回长格式

假设你的原始data.table为dt,核心列包含sex, region, data_type, variable, value(替换为你实际的列名):

library(data.table)

# 转宽格式:将Unemployed和Labour force转为单独列
dt_wide <- dcast(dt, sex + region + data_type ~ variable, value.var = "value")

# 计算失业率
dt_wide[, unemployment_rate := `Unemployed total` / `Labour force total`]

# 转回长格式,把失业率作为新的variable行
dt_final <- melt(dt_wide,
                 id.vars = c("sex", "region", "data_type"),
                 measure.vars = c("Unemployed total", "Labour force total", "unemployment_rate"),
                 variable.name = "variable",
                 value.name = "value")

方法2:长格式下直接分组计算

无需转宽,直接在原长格式数据中分组提取对应值计算:

# 分组计算失业率
rate_dt <- dt[, .(
  value = .SD[variable == "Unemployed total", value] / .SD[variable == "Labour force total", value],
  variable = "unemployment_rate"
), by = .(sex, region, data_type)]

# 合并原数据与失业率数据,得到行形式结果
dt_final <- rbind(dt, rate_dt)

关键提示

  • 确保每个分组(sex+region+data_type)下同时存在"Unemployed total"和"Labour force total",否则会产生NA
  • 若有其他维度列(如时间),需将其加入id.vars(方法1)或by参数(方法2)
  • 可通过round()控制失业率小数位数,例如round(Unemployed total/Labour force total, 4)

内容的提问来源于stack exchange,提问作者John Steed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 16:57:43