基于身高排序拆分R语言数据集:函数实现错误排查与修正
Problem Analysis
Your function has three key issues causing overlapping height ranges:
- Missing sorting step: The function doesn't sort the input data by
heightin descending order. If you pass unsorted data (or forget to pre-sort it), splits become random chunks instead of height-based groups. - Incorrect indexing: The loop uses
limits[i]+1for all subsets, skipping the first row of the sorted data for the first group. - Error-prone split points: Using
as.integer(seq(...))leads to floating-point rounding issues that can create inconsistent group sizes.
Correct Solutions
Below are two robust methods to implement your desired split, supporting any number of subsets:
Option 1: Using dplyr::ntile (Simplest Approach)
This method leverages dplyr's built-in ntile function to evenly split sorted data into groups:
library(dplyr) create_h <- function(data, n_splits = 4) { # Sort data by height descending and assign group labels sorted_data <- data %>% arrange(desc(height)) %>% mutate(group = ntile(desc(height), n_splits)) # Split into a list of subsets h_list <- split(sorted_data, sorted_data$group) # Rename groups to h_1, h_2, etc. names(h_list) <- paste0("h_", seq_along(h_list)) return(h_list) } # Usage: # First filter the target data part_1 <- my_data %>% filter(gender == "Male" & status == "Immigrant") # Generate subsets h_list <- create_h(part_1, n_splits = 4) # Optional: Assign subsets to global environment (use list directly for better practice) list2env(h_list, .GlobalEnv)
Option 2: Manual Index Calculation (Controlled Approach)
If you prefer manual control over split points, this method ensures exact group sizes:
library(dplyr) create_h <- function(data, n_splits = 4) { # Sort data by height descending sorted_data <- data %>% arrange(desc(height)) total_rows <- nrow(sorted_data) # Calculate base group size and remainder base_size <- floor(total_rows / n_splits) remainder <- total_rows %% n_splits # Create split limits (adjust first 'remainder' groups to have one extra row) limits <- c(0, cumsum(rep(base_size, n_splits)) + c(rep(1, remainder), rep(0, n_splits - remainder))) # Generate subsets h_list <- list() for(i in 1:n_splits) { start <- limits[i] + 1 end <- limits[i+1] h_list[[i]] <- sorted_data[start:end, ] } names(h_list) <- paste0("h_", seq_along(h_list)) return(h_list) } # Usage same as Option 1
Verification
Both methods will produce subsets with non-overlapping height ranges, matching your manual results:
# Check ranges for each subset lapply(h_list, function(x) range(x$height))
Content of the question来源于stack exchange,提问作者stats_noob
相关产品推荐
相关产品推荐

