在R语言中按组修正年份数值序列的录入错误
修正年份与"Number of years"序列错误的R方案
核心逻辑
每个ID的Value(即Number of years)应满足Value = 基准值 + (当前年份 - 基准年份),也就是任意两个记录的Value差值 = 年份差值。我们可以通过计算每个点的截距 = Value - Period,找到每个ID中出现频率最高的截距(该截距对应正确的序列规律),再用此截距生成所有年份的正确Value。
实现代码
首先加载数据并排序:
# 加载示例数据 df <- structure(list(id = c(1, 1, 1, 1, 2, 2, 2, 2, 3, 3, 3, 3, 4, 4, 4, 4), Period = c(2005, 2008, 2009, 2010, 2007, 2015, 2016, 2017, 2014, 2015, 2016, 2017, 2007, 2009, 2010, 2011), Value = c(18, 71, 22, 23, 27, 35, 76, 37, 30, 31, 32, 83, 88, 20, 21, 62)), row.names = c(NA, 16L), class = "data.frame") # 按ID和年份排序 library(dplyr) df_sorted <- df %>% arrange(id, Period)
然后计算修正后的Value:
df_corrected <- df_sorted %>% group_by(id) %>% mutate( # 计算每个点的截距 intercept = Value - Period, # 找到当前ID中出现频率最高的截距(即正确序列的截距) correct_intercept = as.numeric(names(which.max(table(intercept)))), # 生成修正后的Value corrected_Value = correct_intercept + Period ) %>% ungroup()
修正结果
执行代码后,df_corrected中的corrected_Value列即为修正后的正确值:
# A tibble: 16 × 6 id Period Value intercept correct_intercept corrected_Value <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> 1 1 2005 18 -1987 -1987 18 2 1 2008 71 -1937 -1987 21 3 1 2009 22 -1987 -1987 22 4 1 2010 23 -1987 -1987 23 5 2 2007 27 -1980 -1980 27 6 2 2015 35 -1980 -1980 35 7 2 2016 76 -1940 -1980 36 8 2 2017 37 -1980 -1980 37 9 3 2014 30 -1984 -1984 30 10 3 2015 31 -1984 -1984 31 11 3 2016 32 -1984 -1984 32 12 3 2017 83 -1934 -1984 33 13 4 2007 88 -1919 -1989 18 14 4 2009 20 -1989 -1989 20 15 4 2010 21 -1989 -1989 21 16 4 2011 62 -1949 -1989 22
补充说明
- 该方法依赖"正确数据的截距出现频率最高"这一前提,若你的数据存在多个截距频率相同的极端情况,可进一步通过连续年份的差值验证来筛选正确截距。
- 若不需要保留原错误值,可直接将
corrected_Value重命名为Value。
内容的提问来源于stack exchange,提问作者Nuñes
相关产品推荐
相关产品推荐

