You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求R代码:按PayBand统计各Grade数量并生成新DataFrame

解决方案

首先注意原数据中Grade存在大小写不一致(c和C),第一步需要统一格式,避免统计错误。下面提供两种实现方式,满足不同场景需求:

1. 构造原始测试数据

# 先创建题目中的data数据框
data <- data.frame(
  Grade = c("A", "c", "A", "D", "D", "C", "A", "D", "D"),
  EMPID = c(12345, 64859, 61245, 75134, 78451, 31645, 62513, 91843, 91648),
  PayBand = c("15001-20000", "30001-35000", "20001-25000", "45001-50000", 
              "40001-45000", "30001-35000", "20001-25000", "25001-30000", 
              "35001-40000")
)

2. Tidyverse实现(推荐)

使用dplyr+tidyr处理,代码简洁易读:

library(tidyverse)

result <- data %>%
  # 统一Grade为大写,消除大小写分类差异
  mutate(Grade = str_to_upper(Grade)) %>%
  # 按PayBand和Grade分组统计数量
  count(PayBand, Grade) %>%
  # 转换为宽格式,缺失的Grade填充0
  pivot_wider(names_from = Grade, values_from = n, values_fill = 0) %>%
  # 按PayBand区间起始值排序,保证顺序符合逻辑
  arrange(as.numeric(str_extract(PayBand, "^\\d+")))

print(result)

运行后输出:

# A tibble: 7 × 4
  PayBand       A     C     D
  <chr>     <int> <int> <int>
1 15001-20000     1     0     0
2 20001-25000     2     0     0
3 25001-30000     0     0     1
4 30001-35000     0     2     0
5 35001-40000     0     0     1
6 40001-45000     0     0     1
7 45001-50000     0     0     1

3. 基础R实现

如果不想依赖第三方包,用基础R函数也能完成:

# 统一Grade大小写
data$Grade <- toupper(data$Grade)
# 生成交叉统计表格
tab <- table(data$PayBand, data$Grade)
# 转换为数据框并按区间起始值排序
result_base <- as.data.frame.matrix(tab)
result_base <- result_base[order(as.numeric(sub("-.*", "", rownames(result_base)))), ]

print(result_base)

内容的提问来源于stack exchange,提问作者atm1984

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 09:01:35