You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pmg回归报错求助:面板数据存在重复(id-time)组合问题

Fixing "duplicate couples (id-time)" Error in pmg (Fama-MacBeth Regression)

Hey there, let's work through this error you're hitting with the pmg function from the plm package. The core issue here is that Fama-MacBeth regression (via pmg) requires your panel data to have unique (time-id) pairs—meaning each fund (id) can only have one observation per month (month). Your dataset currently has duplicate entries for some (month, id) combinations, which breaks the panel structure pmg expects.

Step 1: Identify Exactly Where the Duplicates Are

First, let's pinpoint which (month, id) pairs are duplicated. Follow the hint from the error message:

# Convert your data to a panel data frame
p_fund <- pdata.frame(fund_panel, index = c("month", "id"))

# Create a table to count occurrences of each (month, id) pair
dup_counts <- table(index(p_fund), useNA = "ifany")

# Filter to only show pairs with duplicates
dup_pairs <- dup_counts[dup_counts > 1]
print(dup_pairs)

This will show you exactly which months and fund IDs have repeated rows, so you can verify if it's a data entry error or intended duplicates that need aggregation.

Step 2: Clean the Duplicate Data

Choose a cleaning method based on why the duplicates exist:

Option 1: Remove Duplicates (If They're Data Entry Errors)

If the duplicates are accidental, just keep the first occurrence of each (month, id) pair:

# Base R approach
fund_panel_clean <- fund_panel[!duplicated(fund_panel[, c("month", "id")]), ]

# Or using dplyr for readability
library(dplyr)
fund_panel_clean <- fund_panel %>%
  distinct(month, id, .keep_all = TRUE)

Option 2: Aggregate Duplicates (If Rows Have Different Values)

If the duplicate rows have distinct values (e.g., multiple return entries for the same fund-month), aggregate them using logic that makes sense for your data (e.g., mean, median, sum):

fund_panel_clean <- fund_panel %>%
  group_by(month, id) %>%
  summarise(
    return = mean(return, na.rm = TRUE), # Average return (adjust as needed)
    ex_mkt_ret = mean(ex_mkt_ret, na.rm = TRUE),
    fund_name = first(fund_name), # Keep the first listed fund name
    months = first(months) # Keep the corresponding month string
  ) %>%
  ungroup()

Adjust the aggregation functions (like mean()) to match your analysis needs—for example, use median() if you want to avoid outliers.

Step 3: Re-Run the Fama-MacBeth Regression

Once your data has unique (month, id) pairs, run the pmg function again:

fpmg <- pmg(return ~ ex_mkt_ret, fund_panel_clean, index = c("month", "id"))
summary(fpmg)

A quick note: pmg works fine with unbalanced panels (where some funds are missing in certain months)—the only hard rule is no duplicate (time-id) pairs.

内容的提问来源于stack exchange,提问作者Enrico Dace

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:20:56