如何在R中基于最大长度为数据框补全缺失行
R语言数据框补全:填充NO响应后的缺失行
问题场景
实验数据中,参与者逐词阅读句子并回答YES/NO,一旦回答NO,句子立即结束,后续词的记录缺失。需要补全从NO对应的word_position到max_length之间的所有缺失行,保留participant、sentence_id等静态信息,新行的response设为NO。
解决方案(使用tidyverse工具)
首先加载所需包:
library(tidyverse)
对示例数据进行处理:
test_filled <- test %>% # 按「参与者+句子」分组,确保每组独立处理 group_by(participant, sentence_id) %>% # 计算每组中第一个NO出现的位置,以及句子的最大长度 mutate( first_no_pos = ifelse(any(response == "NO"), min(word_position[response == "NO"]), NA), group_max = max(max_length) ) %>% # 补全该句子所有应存在的word_position行 complete(word_position = 1:group_max) %>% # 填充参与者、句子ID等组内固定信息 fill(participant, sentence_id, group_max, first_no_pos, .direction = "downup") %>% # 设置响应值:原有响应保留,NO位置及之后的新增行设为NO mutate( response = case_when( !is.na(response) ~ response, word_position >= first_no_pos ~ "NO", TRUE ~ NA_character_ ), max_length = group_max ) %>% # 清理临时辅助列,取消分组 select(-first_no_pos, -group_max) %>% ungroup()
补充:填充缺失的word列
如果需要将缺失的word列也补全为句子的正确词汇,可以先建立句子-词位的映射表,再关联到数据中:
# 构建句子完整词序列的映射表 word_lookup <- tribble( ~sentence_id, ~word_position, ~word, "dog_sentence", 1, "the", "dog_sentence", 2, "dog", "dog_sentence", 3, "went", "dog_sentence", 4, "home.", "plant_sentence", 1, "I", "plant_sentence", 2, "watered", "plant_sentence", 3, "my", "plant_sentence", 4, "plants", "plant_sentence", 5, "today." ) # 关联映射表补全word列 test_filled_with_words <- test_filled %>% left_join(word_lookup, by = c("sentence_id", "word_position")) %>% mutate(word = coalesce(word.x, word.y)) %>% select(-word.x, -word.y)
效果验证
查看参与者001的plant_sentence补全结果:
test_filled_with_words %>% filter(participant == "001", sentence_id == "plant_sentence")
输出结果:
| participant | sentence_id | word | word_position | max_length | response |
|---|---|---|---|---|---|
| 001 | plant_sentence | I | 1 | 5 | NO |
| 001 | plant_sentence | watered | 2 | 5 | NO |
| 001 | plant_sentence | my | 3 | 5 | NO |
| 001 | plant_sentence | plants | 4 | 5 | NO |
| 001 | plant_sentence | today. | 5 | 5 | NO |
内容的提问来源于stack exchange,提问作者microcastle
相关产品推荐
相关产品推荐

