You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中基于连续6个月负GDP增长识别衰退期?

识别面板数据中的连续6个月负GDP增长衰退期

针对你的需求,我们可以通过分组计算连续负增长时段的方式来标记衰退期,以下提供两种适合大型面板数据的实现方法(dplyr/tidyverse 和 data.table),后者处理大数据效率更高。

核心逻辑

  1. 按国家分组,确保数据严格按时间排序(这是关键,否则连续时段计算会出错)
  2. 标记每个观测是否为负GDP增长
  3. 对连续的负增长时段进行分组
  4. 判断每个连续负增长分组的长度是否≥6,若是则将该分组内的所有观测标记为crisis=1

方法1:使用dplyr(tidyverse生态)

假设你的数据框名为df,包含country(国家)、month(时间)、gdp_growth(GDP增长率)三列,先模拟一份样本数据方便测试:

library(tidyverse)
set.seed(123)
df <- tibble(
  country = rep(c("A", "B", "C"), each = 24),
  month = rep(1:24, 3),
  gdp_growth = c(rnorm(10, 0.5, 0.1), rep(-0.2, 7), rnorm(7, 0.5, 0.1),
                 rnorm(24, 0.3, 0.1),
                 rep(-0.1, 5), rnorm(4, 0.4, 0.1), rep(-0.3, 6), rnorm(9, 0.5, 0.1))
)

执行以下代码生成crisis变量:

df_crisis <- df %>%
  arrange(country, month) %>%  # 强制按国家+时间排序
  group_by(country) %>%
  mutate(
    is_negative = gdp_growth < 0,
    # 生成连续负增长的分组ID(rleid来自data.table,需先安装data.table包)
    negative_group = rleid(is_negative),
    # 计算每个连续分组的观测数
    group_length = n(),
    # 标记衰退期:负增长且连续时长≥6个月则为1
    crisis = ifelse(is_negative & group_length >= 6, 1, 0)
  ) %>%
  ungroup()

如果不想依赖data.table的rleid,可以用纯dplyr实现连续分组:

df_crisis <- df %>%
  arrange(country, month) %>%
  group_by(country) %>%
  mutate(
    is_negative = gdp_growth < 0,
    # 当当前行的负增长状态与上一行不同时,分组ID+1
    negative_group = cumsum(is_negative != lag(is_negative, default = !is_negative[1])),
    group_length = n(),
    crisis = ifelse(is_negative & group_length >= 6, 1, 0)
  ) %>%
  ungroup()

方法2:使用data.table(适合超大型数据)

data.table的分组计算效率远高于dplyr,处理百万级以上面板数据更友好:

library(data.table)
# 将数据框转为data.table格式
setDT(df)
# 按国家和时间排序
setkey(df, country, month)

# 生成负增长标记、连续分组ID
df[, `:=`(
  is_negative = gdp_growth < 0,
  negative_group = rleid(is_negative)
), by = country]

# 计算每个连续分组的长度,标记衰退期
df[, group_length := .N, by = .(country, negative_group)]
df[, crisis := as.integer(is_negative & group_length >= 6)]

验证结果

你可以通过以下代码检查标记是否正确:

# 筛选出所有衰退期观测,查看连续时长
df_crisis %>%
  filter(crisis == 1) %>%
  group_by(country, negative_group) %>%
  summarize(start_month = min(month), end_month = max(month), duration = n())

注意事项

  • 确保你的时间变量(month或日期列)是连续且有序的,若存在缺失值,需先处理(比如补全缺失时段或标记为非负增长)
  • 如果时间是日期格式,将排序逻辑改为arrange(country, date)或setkey(df, country, date)

内容的提问来源于stack exchange,提问作者CF96

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 19:06:29