R语言按行统计各区间内数值个数的技术实现问询
Hey there! Let's tackle this problem where you need to count, for each row in your R data frame, how many values fall into four specific intervals: <25, 25-49, 50-74, and >=75. Your initial code is actually a great, straightforward solution—let's walk through it first, then look at a couple of alternative approaches for more flexibility.
Your Original Code Works Perfectly
First, let's recap your sample data setup and your counting logic:
Step 1: Generate Sample Data
set.seed(007) x <- data.frame( v1 = sample(1:100, 50), v2 = sample(1:100, 50), v3 = sample(1:100, 50), v4 = sample(1:100, 50), v5 = sample(1:100, 50) )
Step 2: Row-Wise Bin Counting
Your code uses rowSums() to tally values in each bin, and it's totally effective:
# Count values less than 25 per row x$less.25 <- rowSums(x < 25, na.rm = TRUE) # Count values between 25 and 49 (inclusive of 25, exclusive of 50) x$between.25_49 <- rowSums(x >= 25 & x < 50, na.rm = TRUE) # Count values between 50 and 74 (inclusive of 50, exclusive of 75) x$between.50_74 <- rowSums(x >= 50 & x < 75, na.rm = TRUE) # Count values 75 or higher per row x$greater.75 <- rowSums(x >= 75, na.rm = TRUE)
Why this works:
- When you run a comparison like
x < 25, R returns a boolean data frame where each cell isTRUEif the value meets the condition,FALSEotherwise. rowSums()automatically treatsTRUEas 1 andFALSEas 0, so summing across each row gives you the exact count of values in that bin for the row.- The
na.rm = TRUEflag is a smart touch—it ignores any missing values (even though your sample data doesn't have NAs, this is essential for real-world datasets).
Alternative: Scalable Binning with apply() and cut()
If you ever need to adjust your bins or add more columns later, this method is more scalable. You define your bins once, then use apply() to categorize and count values per row:
# Define your bin boundaries and labels bins <- c(-Inf, 25, 50, 75, Inf) bin_labels <- c("less.25", "between.25_49", "between.50_74", "greater.75") # Calculate bin counts for each row row_bin_counts <- t(apply(x, 1, function(row) { # Categorize each value, then tabulate counts table(cut(row, breaks = bins, labels = bin_labels, include.lowest = TRUE)) })) # Combine counts with original data frame x_with_counts <- cbind(x, as.data.frame(row_bin_counts))
This way, you only need to tweak the bins or bin_labels vectors if your interval needs change, instead of writing a separate rowSums() line for each bin.
Alternative: Tidyverse Approach with dplyr & tidyr
If you prefer the tidyverse syntax, you can reshape your data to long format, categorize values, count per row, then reshape back to wide format:
library(dplyr) library(tidyr) # Reuse the bins and bin_labels from the previous example x_with_counts <- x %>% mutate(row_id = row_number()) %>% # Add a unique ID for each row pivot_longer(-row_id, names_to = "column", values_to = "value") %>% # Reshape to long format mutate(bin = cut(value, breaks = bins, labels = bin_labels, include.lowest = TRUE)) %>% # Assign bins count(row_id, bin) %>% # Count values per row and bin pivot_wider(names_from = bin, values_from = n, values_fill = 0) %>% # Reshape back to wide select(-row_id) %>% # Remove the row ID cbind(x, .) # Merge with original data
This method is super readable if you're used to tidyverse tools, and it's easy to extend with other data transformations if needed.
No matter which method you choose, you'll end up with your original data frame plus four new columns, each showing the number of values in the corresponding interval for every row.
内容的提问来源于stack exchange,提问作者Borexino

