You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中统计month_pct和year_pct列的正负零值数量及占比

统计百分比列的正负零数量及占比(R语言实现)

嘿,这个需求很明确,咱们用R一步步来搞定它~首先你的数据集里month_pct和year_pct是因子类型的百分比字符串,没法直接用来判断正负,所以第一步得先把它们转换成可计算的数值型,然后再做统计。

步骤1:加载并预处理数据

先把你的数据集加载进来,然后把百分比因子转换成数值:

# 加载数据集
df <- structure(list(id = 1:11, price = c(40.59, 70.42, 1.8, 1.98, 65.02, 2.23, 54.79, 54.7, 3.32, 1.77, 3.5), month_pct = structure(c(11L, 10L, 9L, 8L, 7L, 6L, 5L, 4L, 3L, 1L, 2L), .Label = c("-19.91%", "-8.55%", "1.22%", "1.39%", "1.41%", "1.83%", "2.02%", "2.59%", "2.86%", "6.58%", "8.53%"), class = "factor"), year_pct = structure(c(4L, 9L, 5L, 3L, 10L, 1L, 11L, 8L, 6L, 7L, 2L), .Label = c("-10.44%", "-19.91%", "-2.46%", "-35.26%", "-4.26%", "-5.95%", "-6.35%", "-6.91%", "-7.95%", "1.51%", "1.54%"), class = "factor")), class = "data.frame", row.names = c(NA, -11L))

# 将因子型百分比转换为数值型(去掉%符号,转数值后除以100)
df$month_pct_num <- as.numeric(gsub("%", "", as.character(df$month_pct))) / 100
df$year_pct_num <- as.numeric(gsub("%", "", as.character(df$year_pct))) / 100

步骤2:编写统计函数

我们可以写一个通用函数,输入一个数值向量,就能返回它的负值、零值、正值的数量和占比:

stat_pct <- function(x) {
  # 统计各类别数量
  neg_count <- sum(x < 0, na.rm = TRUE)
  zero_count <- sum(x == 0, na.rm = TRUE)
  pos_count <- sum(x > 0, na.rm = TRUE)
  total_count <- neg_count + zero_count + pos_count
  
  # 计算占比(保留两位小数)
  neg_pct <- round(neg_count / total_count * 100, 2)
  zero_pct <- round(zero_count / total_count * 100, 2)
  pos_pct <- round(pos_count / total_count * 100, 2)
  
  # 返回结构化结果
  data.frame(
    category = c("负值", "零值", "正值"),
    count = c(neg_count, zero_count, pos_count),
    percentage = paste0(c(neg_pct, zero_pct, pos_pct), "%")
  )
}

步骤3:应用函数并整合结果

把函数分别应用到处理后的两列,再合并结果:

# 统计month_pct的情况
month_stat <- stat_pct(df$month_pct_num)
month_stat$column <- "month_pct"

# 统计year_pct的情况
year_stat <- stat_pct(df$year_pct_num)
year_stat$column <- "year_pct"

# 合并并整理结果
final_result <- rbind(month_stat, year_stat)
final_result <- final_result[, c("column", "category", "count", "percentage")]

# 查看最终结果
print(final_result)

运行结果示例

执行完上面的代码后,你会得到这样的输出:

column category count percentage
1 month_pct      负值     2     18.18%
2 month_pct      零值     0      0.00%
3 month_pct      正值     9     81.82%
4  year_pct      负值     9     81.82%
5  year_pct      零值     0      0.00%
6  year_pct      正值     2     18.18%

可选:用dplyr简化流程(适合熟悉tidyverse的用户)

如果你习惯用tidyverse工具链,可以用管道操作一步完成,不需要新增列:

library(dplyr)
library(tidyr)

df %>%
  mutate(
    across(c(month_pct, year_pct), ~ as.numeric(gsub("%", "", as.character(.x))) / 100, .names = "{.col}_num")
  ) %>%
  summarise(
    across(ends_with("_num"), ~ list(stat_pct(.x)))
  ) %>%
  pivot_longer(everything(), names_to = "column", values_to = "stats") %>%
  mutate(column = gsub("_num", "", column)) %>%
  unnest(stats)

内容的提问来源于stack exchange,提问作者ah bon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 11:42:33