如何在满足条件时为DataFrame新增行?以DNP触发场景为例
修正含'DNP'行的DataFrame错位问题
需求说明
当DataFrame的div列出现'DNP'值时,需要将下一行的错位数据调整为新增行,同时清空'DNP'行的stat列,确保各列数据对应正确字段。
错误与正确数据示例
错误的DataFrame(数据错位)
incorrect_df <- data.frame( game = c(1:3, '2023-11-17'), Date = c('2023-11-20', '2023-11-19', '2023-11-18', 'C'), div = c(LETTERS[1:2], 'DNP', '18'), stat = c(14, 29, 4, '') )
对应表格:
| game | Date | div | stat |
|---|---|---|---|
| 1 | 2023-11-20 | A | 14 |
| 2 | 2023-11-19 | B | 29 |
| 3 | 2023-11-18 | DNP | 4 |
| 2023-11-17 | C | 18 |
目标正确的DataFrame
correct_df <- data.frame( game = c(1:3,4), Date = c('2023-11-20','2023-11-19','2023-11-18','2023-11-17'), div = c(LETTERS[1:2],'DNP',LETTERS[3]), stat = c(14,29,'',18) )
对应表格:
| game | Date | div | stat |
|---|---|---|---|
| 1 | 2023-11-20 | A | 14 |
| 2 | 2023-11-19 | B | 29 |
| 3 | 2023-11-18 | DNP | |
| 4 | 2023-11-17 | C | 18 |
解决方案(R语言实现)
方法1:基础R代码(逻辑直观)
# 复制错误数据 df <- incorrect_df # 定位所有div为'DNP'的行索引 dnp_indices <- which(df$div == 'DNP') # 遍历每个DNP行,处理数据错位 for (idx in dnp_indices) { # 清空当前DNP行的stat值 df$stat[idx] <- '' # 从下一行提取错位数据,整理成正确的新行 new_row <- data.frame( game = as.character(as.integer(df$game[idx]) + 1), Date = df$game[idx + 1], div = df$Date[idx + 1], stat = df$div[idx + 1] ) # 将新行插入到DNP行之后,同时移除原来的错位行 df <- rbind( df[1:idx, ], new_row, df[-(1:(idx + 1)), ] ) } # 可选:将game列转为整数类型 df$game <- as.integer(df$game)
方法2:dplyr简洁实现(适合tidyverse用户)
library(dplyr) df_corrected <- incorrect_df %>% # 标记DNP行及对应的下一行 mutate( is_dnp = div == 'DNP', is_shifted = lead(is_dnp, default = FALSE) ) %>% # 处理DNP行:清空stat mutate(stat = ifelse(is_dnp, '', stat)) %>% # 处理错位行:把值映射到正确的列 mutate( Date = ifelse(is_shifted, game, Date), div = ifelse(is_shifted, Date, div), stat = ifelse(is_shifted, div, stat), game = ifelse(is_shifted, as.character(as.integer(game[is_dnp]) + 1), game) ) %>% # 移除标记列并修正game列类型 select(-is_dnp, -is_shifted) %>% mutate(game = as.integer(game))
说明
基础R代码逻辑清晰,适合新手理解每一步操作;dplyr代码更简洁,可高效处理多组DNP行的批量修正场景。
内容的提问来源于stack exchange,提问作者steel_your_face
相关产品推荐
相关产品推荐

