You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中二元面板数据集矩形化的单行代码方案问询

Solution for Rectangularizing Two-Way Panel Data in R (Like Stata's fillin)

Hey there! I totally get wanting a clean, concise way to expand your panel data to all possible combinations of id_1, id_2, and Year—especially since the old rectangularize function from SciencesPo is no longer usable. Here are two straightforward, efficient solutions that work like Stata's fillin command, with minimal code:

Option 1: Tidyverse (dplyr + tidyr) – Readable & Intuitive

Using the tidyverse ecosystem, you can generate all combinations and join back to your original data in one short pipeline:

library(tidyverse)

# Set seed for reproducible Val values
set.seed(123)
nsamp1 <- data.frame("id_1" = c("A", "A", "B", "C"), 
                     "id_2" = c("a","b", "a","b"), 
                     "Year" = c(1990, 1991, 1990, 1991), 
                     "Val" = runif(4, min=0, max=100), 
                     stringsAsFactors = FALSE)

# Short pipeline to rectangularize
rectangularized_data <- nsamp1 %>%
  expand(id_1, id_2, Year) %>%
  left_join(nsamp1, by = c("id_1", "id_2", "Year"))
  • expand(id_1, id_2, Year) creates every possible Cartesian product of your three grouping variables.
  • left_join() matches the original Val values to their corresponding combinations, filling missing ones with NA exactly like you need.

Option 2: data.table – Ultra-Fast for Large Datasets

If you're working with big data, data.table offers a lightning-fast, one-line solution:

library(data.table)

set.seed(123)
nsamp1 <- data.frame("id_1" = c("A", "A", "B", "C"), 
                     "id_2" = c("a","b", "a","b"), 
                     "Year" = c(1990, 1991, 1990, 1991), 
                     "Val" = runif(4, min=0, max=100), 
                     stringsAsFactors = FALSE)

# Convert to data.table and rectangularize in one step
setDT(nsamp1)
rectangularized_data <- nsamp1[CJ(id_1, id_2, Year, unique = TRUE), on = .(id_1, id_2, Year)]
  • CJ(..., unique=TRUE) generates all unique combinations of your grouping variables.
  • The bracket syntax x[y, on=...] performs a join that automatically fills missing Val entries with NA.

Both methods will produce exactly the output format you showed, no messy multi-step workarounds needed. Just pick the one that fits your workflow best!

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 09:14:09