You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中统计末尾满足条件的连续行数的优化需求

Optimized Solution for Counting Trailing Consecutive Rows Meeting a Condition

Your nested loop approach works but is inefficient, especially for large data frames, because explicit loops in R are slow compared to vectorized operations. Here are two faster, vectorized approaches to solve this problem:

Approach 1: Using Vectorized Checks with which()

This method reverses each column, finds the first position where the value fails the x > 3 condition, and calculates the count of consecutive valid rows before that point.

# Sample data frame
df <- data.frame(
  a = c(1,4,3,6,5,3,5,6,2,6,7,1,4,5,7),
  b = c(1,5,5,4,8,5,1,4,8,1,5,4,8,7,6),
  c = c(4,5,5,8,6,1,4,8,5,5,1,5,1,4,7)
)

# Function to count trailing consecutive valid rows
count_trailing_valid <- function(v) {
  reversed <- rev(v)
  # Find first position where value is <=3
  first_invalid <- which(reversed <= 3)[1]
  # If all values are valid, return total rows; else return position minus 1
  if (is.na(first_invalid)) length(v) else first_invalid - 1
}

# Apply function to each column
results <- sapply(df, count_trailing_valid)

# Output as named vector (convert to data frame if needed)
results
#> a b c 
#> 3 5 2

# To get a data frame output:
results_df <- as.data.frame(t(results))

Approach 2: Using Run-Length Encoding (rle())

This method uses run-length encoding on the reversed column to directly get the length of the first valid run (since reversing turns trailing rows into leading rows):

count_trailing_valid_rle <- function(v) {
  run_info <- rle(rev(v) > 3)
  # If first run is valid (TRUE), return its length; else return 0
  if (run_info$values[1]) run_info$lengths[1] else 0
}

# Apply to columns
results_rle <- sapply(df, count_trailing_valid_rle)
results_rle
#> a b c 
#> 3 5 2

Why These Are Faster

Both approaches use vectorized operations implemented in C under the hood, which are orders of magnitude faster than explicit loops in R. They avoid iterating row-by-row for each column, making them scalable even for large data frames with thousands of rows/columns.

Content of the question originates from Stack Exchange, question author: Vinicius B. de S. Moreira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 22:56:29