You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算年度申请人增长率?附2020-2022数据集求助

问题描述

我有一个包含20个变量的数据集,时间范围为2020-2022年,希望计算年度申请人增长率。已尝试对数据进行subset(子集化)操作,但后续思路受阻。核心需求是将申请人对应到所属年份,进而计算增长率。数据集示例如下:

Observations ID#   Date
 1           1226  2022-10-16
 2           1225  2021-10-15
 3           1224  2020-08-14
 4           1223  2021-12-02
 5           1222  2022-02-25
解决方案

核心逻辑

  1. 从Date字段提取年份,将每个申请人匹配到对应年度
  2. 统计每年的申请人总数(以唯一ID#为计数依据)
  3. 按公式计算增长率:(当年申请人数量 - 上年申请人数量) / 上年申请人数量 * 100%

用R实现

  1. 提取年份并统计年度申请人数量
    假设数据集名为df,先处理日期格式并提取年份,再按年度计数:
# 转换日期格式
df$Date <- as.Date(df$Date)
# 提取年份列
df$Year <- format(df$Date, "%Y")
# 统计每年的唯一申请人数量
yearly_counts <- table(df$Year)
# 转为数据框方便后续计算
yearly_counts_df <- as.data.frame(yearly_counts)
colnames(yearly_counts_df) <- c("Year", "Applicant_Count")
  1. 计算年度增长率
    借助dplyr包的lag()函数获取上年数据,计算增长率:
library(dplyr)
yearly_growth <- yearly_counts_df %>%
  arrange(Year) %>%
  mutate(Growth_Rate = (Applicant_Count - lag(Applicant_Count))/lag(Applicant_Count)*100)

结果中Growth_Rate列即为年度增长率,2020年因无上年数据会显示NA。


用Python实现

  1. 提取年份并统计年度申请人数量
    使用pandas库处理,数据集名为df:
import pandas as pd
# 转换日期格式
df['Date'] = pd.to_datetime(df['Date'])
# 提取年份列
df['Year'] = df['Date'].dt.year
# 按年份统计唯一申请人数量
yearly_counts = df.groupby('Year')['ID#'].nunique().reset_index(name='Applicant_Count')
  1. 计算年度增长率
    用shift()函数获取上年数据,计算增长率:
yearly_counts = yearly_counts.sort_values('Year')
yearly_counts['Growth_Rate'] = (yearly_counts['Applicant_Count'] - yearly_counts['Applicant_Count'].shift(1)) / yearly_counts['Applicant_Count'].shift(1) * 100

2020年的增长率会显示NaN,可根据需求选择填充或保留。

内容的提问来源于stack exchange,提问作者Ram

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 03:46:12