使用R将含起止年份的数据框转换为完整时间序列
Hey there! Expanding your data frame to include every year between start_year and end_year (while keeping all other variable values intact) is a common task, and it's easy to pull off with tidyverse tools. Let's walk through this step by step using your example data.
第一步:补全并确认原始数据
First, let's finish the sample data you provided (I filled in plausible end_year values for demonstration):
data_original <- data.frame( name = c("peter", "peter", "eric", "denisse"), lastname = c("smith", "smith", "jordan", "williams"), age = c(54, 54, 48, 40), start_year = c(1980, 1986, 1990, 2000), end_year = c(1985, 1990, 1995, 2005) )
第二步:用tidyverse实现时间序列展开
We'll use dplyr for row-wise processing and tidyr to unnest the generated year sequences. This is the most readable and scalable approach:
# 加载必要的包 library(tidyverse) # 展开时间序列 data_expanded <- data_original %>% # 按行处理每个时间段 rowwise() %>% # 生成当前行起止年份之间的所有年份(存为列表) mutate(year = list(seq(start_year, end_year))) %>% # 将列表中的年份拆分为单独的行,保留其他变量值 unnest(year) %>% # 可选:移除原始的起止年份列(按需保留) select(-start_year, -end_year) %>% # 按姓名和年份排序,让结果更整洁 arrange(name, lastname, year) # 查看结果 data_expanded
代码解释
rowwise(): Ensures each row is processed independently (so we generate a unique year sequence for everystart_year/end_yearpair).mutate(year = list(seq(start_year, end_year))): Creates a list column containing all years between the start and end for each row.unnest(year): Expands the list column into individual rows, copying the other variables (name, lastname, age) for each year.select(): Optional cleanup to remove the original start/end columns if you don't need them anymore.arrange(): Sorts the final data frame for readability.
输出结果
The resulting data frame will look like this (truncated for brevity):
# A tibble: 22 × 4 name lastname age year <chr> <chr> <dbl> <int> 1 denisse williams 40 2000 2 denisse williams 40 2001 3 denisse williams 40 2002 4 denisse williams 40 2003 5 denisse williams 40 2004 6 denisse williams 40 2005 7 eric jordan 48 1990 8 eric jordan 48 1991 9 eric jordan 48 1992 10 eric jordan 48 1993 # … with 12 more rows
备选方案:Base R实现
If you prefer not to use tidyverse, here's a base R approach that achieves the same result:
# Base R 方法 data_expanded_base <- do.call(rbind, apply(data_original, 1, function(row) { # 生成当前行的年份序列 years <- seq(as.integer(row["start_year"]), as.integer(row["end_year"])) # 为每个年份创建一行数据 data.frame( name = row["name"], lastname = row["lastname"], age = as.integer(row["age"]), year = years, stringsAsFactors = FALSE ) })) # 排序结果 data_expanded_base <- data_expanded_base[order(data_expanded_base$name, data_expanded_base$lastname, data_expanded_base$year), ]
额外提示
- If your
agevariable should increase each year (instead of staying fixed), you can adjust the code to calculate it dynamically. For example, replace theunneststep with:unnest(year) %>% mutate(age = age + (year - start_year)) - Make sure
start_yearandend_yearare integer types (not characters) to avoid errors withseq().
内容的提问来源于stack exchange,提问作者Victoria Bread

