You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R的dplyr在原有计算中新增水仙球茎年月频次列

问题背景

我有一个388×729的大型数据框,需要计算2012-2024这14年间,每年每月的新水仙球茎(数值型列)频次。目前已成功算出该实验每年每月的开展天数,希望在现有代码生成的数据框中新增一列展示水仙球茎的年月频次,同时想知道能否一次性完成这两项计算。

数据框结构

$Year               : num  2012 2012 2012 2012 2012 ... 
$ Month             : Factor w/ 18 levels "April","April ",..: 9 8 8 8 8 8 8 8 8 1 ...
$ Daffodil Bulbs    : num  0 3 0 3 2 1 0 0 0 0 ...

现有计算开展天数的代码

Frequency_Days_Experiment <- MyDf %>% 
  mutate(Month = factor(trimws(Month), levels = month.name, ordered = TRUE)) %>% 
  group_by(Year, Month) %>% 
  count()

模拟数据代码

tibble(
  Month = sample(month.name, 120, replace = TRUE),
  Year = sample(2012:2024, 120, replace = TRUE),
  Number_Daffodils = sample(1:5, 120, replace = TRUE)
) 

当前代码输出

# A tibble: 96 × 3
# Groups:   Year, Month [96]
    Year Month        n
   <dbl> <ord>    <int>
 1  2012 January      1
 2  2012 February     8
 3  2012 April       18
 4  2012 May         21
 5  2012 June        27
 6  2012 July        12
 7  2012 October     12
 8  2012 November     4
 9  2012 December     3
10  2013 February     2
# ℹ 86 more rows

期望输出

# A tibble: 96 × 4
# Groups:   Year, Month [96]
    Year Month        n Frequency_Daffodils
   <dbl> <ord>    <int>                <int>
 1  2012 January      1                   25
 2  2012 February     8                    5
 3  2012 April       18                   13 
 4  2012 May         21                   45
 5  2012 June        27                    9
 6  2012 July        12                   78
 7  2012 October     12                   12
 8  2012 November     4                   62
 9  2012 December     3                    1
10  2013 February     2                    8
# ℹ 86 more rows

解决方案

可以一次性完成两项计算,只需在分组后用summarize同时统计两个指标即可:

完整代码

针对模拟数据(Number_Daffodils列)

result <- MyDf %>% 
  mutate(Month = factor(trimws(Month), levels = month.name, ordered = TRUE)) %>% 
  group_by(Year, Month) %>% 
  summarize(
    Days_Experiment = n(),  # 统计实验开展天数
    Frequency_Daffodils = sum(Number_Daffodils, na.rm = TRUE),  # 统计水仙球茎总频次
    .groups = "drop_last"
  )

针对原始数据(Daffodil Bulbs列)

result <- MyDf %>% 
  mutate(Month = factor(trimws(Month), levels = month.name, ordered = TRUE)) %>% 
  group_by(Year, Month) %>% 
  summarize(
    Days_Experiment = n(),
    Frequency_Daffodils = sum(`Daffodil Bulbs`, na.rm = TRUE),
    .groups = "drop_last"
  )

代码说明

  • trimws(Month):去除月份文本首尾空格,避免因空格导致的因子水平不一致(比如原始数据中的"April "和"April")。
  • factor(..., levels = month.name, ordered = TRUE):将月份转为有序因子,保证结果按自然月份顺序排列。
  • summarize:分组后同时计算两个核心指标:
    • Days_Experiment = n():统计每组行数,即当月实验开展天数。
    • Frequency_Daffodils = sum(..., na.rm = TRUE):对水仙球茎数值列求和,na.rm = TRUE用于忽略数据中可能存在的缺失值。
  • .groups = "drop_last":保留年份分组、取消月份分组,也可根据需求改为"drop"(完全取消分组)或"keep"(保留所有分组)。

内容的提问来源于stack exchange,提问作者Alice Hobbs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 16:22:42