You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于tidyverse/dplyr计算两个R数据框(Aug与Sept)的差异统计量(数量及占比)

嘿,我来帮你用tidyverse工具搞定这个数据对比的需求!先把数据构建出来,然后一步步实现三个计算目标:

首先,我们先加载tidyverse包并创建题目中的两个数据框:

library(tidyverse)

# 构建Sept数据框(9887到9900的连续数值)
Sept <- tibble(value = 9887:9900)

# 构建Aug数据框:前5个是9887-9891,然后3223重复6次,6563重复5次
Aug <- tibble(value = c(9887:9891, rep(3223, 6), rep(6563, 5)))

接下来咱们逐个完成需求:


1. 计算Aug中未出现在Sept中的数值的数量及占比

这里分两种常见统计维度:包含重复值的总观测数和去重后的唯一值,我都给你写出来:

统计总观测的情况:

aug_not_in_sept <- Aug %>%
  # 筛选出Aug里不在Sept中的观测
  anti_join(Sept, by = "value") %>%
  # 计算数量和占比
  summarise(
    count = n(),
    total_aug_observations = nrow(Aug),
    proportion = count / total_aug_observations
  )

# 查看结果
aug_not_in_sept

运行后会得到:Aug里有11个观测不在Sept中,占总观测数的68.75%。

统计唯一值的情况:

aug_unique_not_in_sept <- Aug %>%
  # 先对Aug的数值去重
  distinct(value) %>%
  anti_join(Sept, by = "value") %>%
  summarise(
    unique_count = n(),
    total_unique_aug_values = n_distinct(Aug$value),
    unique_proportion = unique_count / total_unique_aug_values
  )

aug_unique_not_in_sept

结果是:Aug里有2个唯一值(3223和6563)不在Sept中,占Aug唯一值总数的约28.57%。


2. 计算Sept中未出现在Aug中的数值的数量及占比

因为Sept里的数值都是唯一的,所以总观测和唯一值的统计结果一致:

sept_not_in_aug <- Sept %>%
  anti_join(Aug, by = "value") %>%
  summarise(
    count = n(),
    total_sept_observations = nrow(Sept),
    proportion = count / total_sept_observations
  )

sept_not_in_aug

运行结果:Sept里有9个观测(9892到9900)不在Aug中,占总观测数的约64.29%。


3. 计算Aug与Sept共有的数值的数量及占比

同样分两种维度统计:

共有观测数及各自占比:

# 先统计共有观测的数量
common_obs_count <- inner_join(Aug, Sept, by = "value") %>% nrow()

# 计算在两个数据框中的占比
common_stats <- tibble(
  common_observation_count = common_obs_count,
  aug_common_proportion = common_obs_count / nrow(Aug),
  sept_common_proportion = common_obs_count / nrow(Sept)
)

common_stats

结果显示:两者共有5个观测(9887到9891),占Aug总观测数的31.25%,占Sept总观测数的约35.71%。

共有唯一值数及各自占比:

common_unique_count <- inner_join(distinct(Aug), distinct(Sept), by = "value") %>% nrow()

common_unique_stats <- tibble(
  common_unique_value_count = common_unique_count,
  aug_unique_common_proportion = common_unique_count / n_distinct(Aug$value),
  sept_unique_common_proportion = common_unique_count / n_distinct(Sept$value)
)

common_unique_stats

结果是:两者共有5个唯一值,占Aug唯一值总数的约71.43%,占Sept唯一值总数的约35.71%。


内容的提问来源于stack exchange,提问作者Yomi.blaze93

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 18:04:04