基于年份创建时间序列变量t的R代码问题及优化咨询
t in Your R Data Frame Hey there! Let's break down why your code isn't running and walk through way more efficient ways to create that t variable you need.
Why Your Original Code Failed
The core issue is a classic R beginner mistake: you used the assignment operator = when you meant to use the equality check operator ==.
In your code line data[data$year = 2005, "t"] <- 1, you're trying to assign the value 2005 to data$year (which is not what you want) instead of checking which rows have a year equal to 2005. That's why R throws an error—this assignment doesn't make sense in a subsetting context.
Efficient Ways to Create Your t Variable
You don't need to manually assign values for each year! Here are three straightforward methods, depending on your needs:
Method 1: Simple Math (Fastest for Your Exact Case)
Since your years are a continuous sequence (2005 → 1, 2006 → 2, ..., 2010 → 6), you can just calculate t directly with basic arithmetic:
# Add the t column to your data frame data$t <- data$year - 2004
This works because 2005 - 2004 = 1, 2006 - 2004 = 2, and so on—perfect match for your desired output.
Method 2: Tidyverse Style with dplyr
If you prefer working with the tidyverse ecosystem, use mutate() to create the new column in a data pipeline:
library(dplyr) data <- data %>% mutate(t = year - 2004)
This is great if you're already doing other data cleaning/manipulation with tidyverse tools.
Method 3: Generalized for Non-Continuous Years
If your years ever end up non-continuous (or you want a method that works regardless of start year), convert the year to a factor and then to an integer:
# First, sort by gvkey and year to ensure correct ordering (optional but safe) data <- data[order(data$gvkey, data$year), ] # Assign t as the integer representation of the year factor data$t <- as.integer(factor(data$year, levels = unique(data$year)))
This will assign 1 to the earliest year in your data, 2 to the next, and so on—no matter what the actual year values are.
Example Output
Let's test Method 1 with your sample data:
# Create sample data data <- data.frame( gvkey = c(1004, 1004, 1013), year = c(2005, 2006, 2010) ) # Generate t data$t <- data$year - 2004 # Print result data
Output:
gvkey year t 1 1004 2005 1 2 1004 2006 2 3 1013 2010 6
Exactly what you were aiming for!
内容的提问来源于stack exchange,提问作者Yaron Nolan

