R语言中如何将特定行的列数据移至相邻行对应列并修正数据框
修复数据框中错位的评论与数值字段
原始数据框
df <- data.frame(col1=c('71711', '71711', '71711', 'Comment 4', '71711', 'Comment 6'), col2=c('Comment 1','Comment 2','Comment 3', '24','Comment 5','26'), col3 = c('21','22','23',NA,'25',NA), stringsAsFactors = FALSE) # 避免因子类型干扰后续操作
期望输出
Col1 Col2 Col3 71711 Comment 1 21 71711 Comment 2 22 71711 Comment 3 23 71711 Comment 4 24 71711 Comment 5 25 71711 Comment 6 26
你的代码问题分析
你写的循环存在几个关键错误:
- 条件判断逻辑错误:
any(sapply(df$col2, is.numeric)) == "True"完全不成立。因为col2是字符类型,sapply(df$col2, is.numeric)会全部返回FALSE,any()结果也是FALSE;而且布尔值不能和字符串"True"比较,应该直接用TRUE。另外这个条件是检查整个列,不是当前行的情况。 - 代码块缩进错误:
if语句后没有用{}包裹两行赋值代码,导致df$col1[i] <- df$col2[i]不管条件是否满足都会执行,直接破坏了所有行的col1值。 - 逻辑定位错误:没有正确识别错位的行——错位行的特征是
col1是评论、col2是数字、col3为NA,需要针对性处理这些行。
正确解决方法
方法1:基础R实现
直接定位错位的行(col3为NA的行),交换对应字段的值:
# 定位col3为NA的行 na_rows <- is.na(df$col3) # 把错位行的col2数值移到col3 df$col3[na_rows] <- df$col2[na_rows] # 把错位行的col1评论移到col2 df$col2[na_rows] <- df$col1[na_rows] # 把错位行的col1统一设为'71711' df$col1[na_rows] <- '71711' # 查看结果 print(df)
执行后输出:
col1 col2 col3 1 71711 Comment 1 21 2 71711 Comment 2 22 3 71711 Comment 3 23 4 71711 Comment 4 24 5 71711 Comment 5 25 6 71711 Comment 6 26
方法2:dplyr简洁实现
如果你习惯用tidyverse系列工具,可以用dplyr更直观地处理:
library(dplyr) library(stringr) df_fixed <- df %>% mutate( # 标记错位行:col1是评论开头、col2是数字、col3为NA is_misplaced = str_detect(col1, "^Comment") & str_detect(col2, "^\\d+$") & is.na(col3), # 交换错位行的字段值 col3 = ifelse(is_misplaced, col2, col3), col2 = ifelse(is_misplaced, col1, col2), # 统一col1为'71711' col1 = '71711' ) %>% select(-is_misplaced) # 移除标记列 print(df_fixed)
这个方法更灵活,即使错位行的特征有变化,也可以调整is_misplaced的判断条件。
内容的提问来源于stack exchange,提问作者Bubbles
相关产品推荐
相关产品推荐

