You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中统计数据框从首列开始的连续相同值个数

问题描述

我在R语言中有如下结构的数据框:

2021    2020    2019    2018    2017    2015    2010    2006    2002    1998    1994    1990
1       6       6       6       6       4       6       6       6       6       6       6       6
2       6       6       6       6       6       6       6       6       6       3       4       4
3       7       6       6       6       4       6       6       6       6       6       6       6
4       6       6       6       6       6       6       6       6       6       6       6       6
5       4       4       7       6       4       6       6       6       6       6       4       6
6       6       6       6       6       6       4       6       6       6       2       6       6
...

我需要统计从第一列开始(包含第一列)的连续相同值的个数,并将结果存入名为age的新数据框。期望输出如下:

age
1       4
2       9
3       1
4       12
5       2
6       5
...

我尝试使用rle函数,但没能成功实现需求,请问有什么解决办法?

补充可复现数据:

df <- structure(list(LC_2021 = c(6, 6, 6, 6, 6, 6), LC_2020 = c(6, 
6, 6, 6, 4, 6), LC_2019 = c(6, 6, 6, 6, 6, 6), LC_2018 = c(6, 
6, 6, 6, 6, 6), LC_2017 = c(6, 6, 6, 6, 6, 6), LC_2015 = c(6, 
6, 6, 6, 6, 6), LC_2010 = c(6, 6, 6, 6, 6, 6), LC_2006 = c(4, 
4, 4, 6, 4, 4), LC_2002 = c(6, 6, 6, 6, 6, 6), LC_1998 = c(6, 
6, 6, 6, 6, 6), LC_1994 = c(6, 6, 6, 6, 6, 6), LC_1990 = c(6, 
6, 6, 6, 6, 6)), row.names = c(NA, 6L), class = "data.frame")
解决方案

方法1:apply结合rle逐行处理

这是最直观的实现方式,对每行数据用rle提取连续值的长度,直接取第一个连续段的长度即可:

# 逐行计算从第一列开始的连续相同值个数
age <- data.frame(age = apply(df, 1, function(x) {
  rle_result <- rle(x)
  rle_result$lengths[1]
}))

# 查看结果
age

运行后输出:

age
1   7
2   7
3   7
4  12
5   1
6   7

方法2:向量化操作(大数据量更高效)

如果数据行数很多,向量化操作比apply的循环更快。通过判断每行中第一个与首列值不同的位置,间接计算连续长度:

# 转换为矩阵方便操作
mat <- as.matrix(df)
# 找到每行中第一个不等于首列值的位置
first_diff <- apply(mat != mat[,1], 1, which.max)
# 处理全列值都相同的行(which.max会返回0)
age <- data.frame(age = ifelse(first_diff == 0, ncol(mat), first_diff - 1))

age

方法3:dplyr行操作实现

如果你习惯使用tidyverse工具链,可以用rowwise结合c_across完成行级计算:

library(dplyr)

age <- df %>%
  rowwise() %>%
  mutate(age = {
    x <- c_across(everything())
    rle(x)$lengths[1]
  }) %>%
  select(age)

age

内容的提问来源于stack exchange,提问作者Xandian97

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 23:44:56