pmg回归报错求助:面板数据存在重复(id-time)组合问题
pmg (Fama-MacBeth Regression) Hey there, let's work through this error you're hitting with the pmg function from the plm package. The core issue here is that Fama-MacBeth regression (via pmg) requires your panel data to have unique (time-id) pairs—meaning each fund (id) can only have one observation per month (month). Your dataset currently has duplicate entries for some (month, id) combinations, which breaks the panel structure pmg expects.
Step 1: Identify Exactly Where the Duplicates Are
First, let's pinpoint which (month, id) pairs are duplicated. Follow the hint from the error message:
# Convert your data to a panel data frame p_fund <- pdata.frame(fund_panel, index = c("month", "id")) # Create a table to count occurrences of each (month, id) pair dup_counts <- table(index(p_fund), useNA = "ifany") # Filter to only show pairs with duplicates dup_pairs <- dup_counts[dup_counts > 1] print(dup_pairs)
This will show you exactly which months and fund IDs have repeated rows, so you can verify if it's a data entry error or intended duplicates that need aggregation.
Step 2: Clean the Duplicate Data
Choose a cleaning method based on why the duplicates exist:
Option 1: Remove Duplicates (If They're Data Entry Errors)
If the duplicates are accidental, just keep the first occurrence of each (month, id) pair:
# Base R approach fund_panel_clean <- fund_panel[!duplicated(fund_panel[, c("month", "id")]), ] # Or using dplyr for readability library(dplyr) fund_panel_clean <- fund_panel %>% distinct(month, id, .keep_all = TRUE)
Option 2: Aggregate Duplicates (If Rows Have Different Values)
If the duplicate rows have distinct values (e.g., multiple return entries for the same fund-month), aggregate them using logic that makes sense for your data (e.g., mean, median, sum):
fund_panel_clean <- fund_panel %>% group_by(month, id) %>% summarise( return = mean(return, na.rm = TRUE), # Average return (adjust as needed) ex_mkt_ret = mean(ex_mkt_ret, na.rm = TRUE), fund_name = first(fund_name), # Keep the first listed fund name months = first(months) # Keep the corresponding month string ) %>% ungroup()
Adjust the aggregation functions (like mean()) to match your analysis needs—for example, use median() if you want to avoid outliers.
Step 3: Re-Run the Fama-MacBeth Regression
Once your data has unique (month, id) pairs, run the pmg function again:
fpmg <- pmg(return ~ ex_mkt_ret, fund_panel_clean, index = c("month", "id")) summary(fpmg)
A quick note: pmg works fine with unbalanced panels (where some funds are missing in certain months)—the only hard rule is no duplicate (time-id) pairs.
内容的提问来源于stack exchange,提问作者Enrico Dace

