You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R语言PLM包做面板数据固定效应回归:重复id-time问题求助

解决国际贸易引力模型面板数据的重复(id-time)警告问题

在使用R的PLM包处理双边贸易面板数据(结构为出口国-贸易伙伴-年份)时,出现duplicate couples (id-time)警告且无法直接删除观测值,可按以下步骤解决:

1. 定位重复的观测组合

先创建面板对象(即使有警告),再用PLM提示的方法找出具体重复的(出口国-贸易伙伴-年份)组合:

# 创建面板对象(保留警告)
panel_data <- pdata.frame(df, index = c("CountryName", "Counterpart_Country_Name", "Year"))
# 生成重复组合的统计表格
dup_table <- table(index(panel_data), useNA = "ifany")
# 筛选出重复次数>1的组合
dup_combinations <- dup_table[dup_table > 1]
print(dup_combinations)

这一步能精准定位哪些双边年份组合存在重复行,方便后续针对性处理。

2. 处理重复观测

根据重复原因选择对应方式:

情况1:数据录入重复

直接基于出口国-贸易伙伴-年份组合去重,保留唯一行:

library(dplyr)
# 去重,保留每组第一行(如需保留最后一行,将.first改为.last)
df_clean <- df %>%
  distinct(CountryName, Counterpart_Country_Name, Year, .keep_all = TRUE)

情况2:细分维度未聚合(如按产品拆分的贸易数据)

引力模型中通常需要将同一双边年份的贸易数据聚合,贸易额取总和,国家层面变量(如人口、GDP)取唯一值:

library(dplyr)
df_clean <- df %>%
  group_by(CountryName, Counterpart_Country_Name, Year) %>%
  summarise(
    Trade = sum(Trade),  # 贸易额求和
    Population = first(Population),  # 同一年份国家人口唯一,取任意值即可
    Distance = first(Distance),
    GDP = first(GDP),
    # 其他变量按性质选择聚合方式,如log变量可先聚合原始值再取对数
    .groups = "drop"
  )
# 重新生成对数变量(如果原始变量聚合后需要)
df_clean$log_pop <- log(df_clean$Population)
df_clean$log_dist <- log(df_clean$Distance)
df_clean$log_GDP <- log(df_clean$GDP)
df_clean$log_Trade <- log(df_clean$Trade)

3. 重新构建面板并执行固定效应回归

处理完重复后,重新创建面板对象并运行回归:

# 创建无重复的面板数据
panel_data_clean <- pdata.frame(df_clean, index = c("CountryName", "Counterpart_Country_Name", "Year"))
# 双向固定效应回归(控制出口国和进口国异质性)
fe_model <- plm(log_Trade ~ log_GDP + log_pop + log_dist, 
                data = panel_data_clean, 
                model = "within",
                effect = "twoways")
# 查看回归结果
summary(fe_model)

如果需要额外控制年份固定效应,可在公式中加入+ factor(Year)。

内容的提问来源于stack exchange,提问作者0klahoma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 10:12:23