R语言:如何获取df2每行在df1中的对应行索引?
获取df2每行在df1中的对应索引
我有两个dataframe:df1和df2,其中df2由df1中的行组成。想要获取df2每行在df1中的对应索引。已知数据中不会出现df2的行在df1中有重复的情况,无需处理该场景。
示例数据
df1 <- data.frame(animal=c('koala', 'hedgehog', 'sloth', 'panda'), country=c('Australia', 'Italy', 'Peru', 'China'), avg_sleep_hours=c(21, 18, 17, 10)) df2 <- data.frame(animal=c('koala', 'sloth', 'panda', 'panda'), country=c('Australia', 'Peru', 'China', 'China'), avg_sleep_hours=c(21,17,10,10))
期望结果
1 3 4 4
我自己的可行代码
findIdxRow <- function(row, df) { n <- nrow(df) is_equal <- sapply(1:n, function(i) all(row==df[i,])) return(which(is_equal)) } indexes <- sapply(1:nrow(df2), function(i) findIdxRow(df2[i,],df1))
更简洁的写法
方法1:Base R 快速实现(interaction + match)
利用interaction将每行的所有列组合成唯一标识,再通过match匹配位置,代码极其简洁:
indexes <- match(interaction(df2), interaction(df1))
方法2:Base R 基于merge实现
给df1添加行索引列后,通过merge按所有列匹配提取索引:
df1_with_idx <- cbind(df1, idx = seq(nrow(df1))) indexes <- merge(df2, df1_with_idx, by = names(df1))$idx
方法3:dplyr 写法(tidyverse风格)
如果习惯使用tidyverse工具链,可以用dplyr的链式操作实现:
library(dplyr) indexes <- df2 %>% left_join(df1 %>% mutate(idx = row_number()), by = names(df1)) %>% pull(idx)
这些方法都避免了逐行循环的冗余操作,代码更简洁且效率更高,尤其适合处理大规模数据。
内容的提问来源于stack exchange,提问作者Ccile
相关产品推荐
相关产品推荐

