You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于身高排序拆分R语言数据集:函数实现错误排查与修正

Problem Analysis

Your function has three key issues causing overlapping height ranges:

  1. Missing sorting step: The function doesn't sort the input data by height in descending order. If you pass unsorted data (or forget to pre-sort it), splits become random chunks instead of height-based groups.
  2. Incorrect indexing: The loop uses limits[i]+1 for all subsets, skipping the first row of the sorted data for the first group.
  3. Error-prone split points: Using as.integer(seq(...)) leads to floating-point rounding issues that can create inconsistent group sizes.
Correct Solutions

Below are two robust methods to implement your desired split, supporting any number of subsets:

Option 1: Using dplyr::ntile (Simplest Approach)

This method leverages dplyr's built-in ntile function to evenly split sorted data into groups:

library(dplyr)

create_h <- function(data, n_splits = 4) {
  # Sort data by height descending and assign group labels
  sorted_data <- data %>%
    arrange(desc(height)) %>%
    mutate(group = ntile(desc(height), n_splits))
  
  # Split into a list of subsets
  h_list <- split(sorted_data, sorted_data$group)
  
  # Rename groups to h_1, h_2, etc.
  names(h_list) <- paste0("h_", seq_along(h_list))
  
  return(h_list)
}

# Usage:
# First filter the target data
part_1 <- my_data %>% filter(gender == "Male" & status == "Immigrant")

# Generate subsets
h_list <- create_h(part_1, n_splits = 4)

# Optional: Assign subsets to global environment (use list directly for better practice)
list2env(h_list, .GlobalEnv)

Option 2: Manual Index Calculation (Controlled Approach)

If you prefer manual control over split points, this method ensures exact group sizes:

library(dplyr)

create_h <- function(data, n_splits = 4) {
  # Sort data by height descending
  sorted_data <- data %>% arrange(desc(height))
  total_rows <- nrow(sorted_data)
  
  # Calculate base group size and remainder
  base_size <- floor(total_rows / n_splits)
  remainder <- total_rows %% n_splits
  
  # Create split limits (adjust first 'remainder' groups to have one extra row)
  limits <- c(0, cumsum(rep(base_size, n_splits)) + c(rep(1, remainder), rep(0, n_splits - remainder)))
  
  # Generate subsets
  h_list <- list()
  for(i in 1:n_splits) {
    start <- limits[i] + 1
    end <- limits[i+1]
    h_list[[i]] <- sorted_data[start:end, ]
  }
  
  names(h_list) <- paste0("h_", seq_along(h_list))
  return(h_list)
}

# Usage same as Option 1
Verification

Both methods will produce subsets with non-overlapping height ranges, matching your manual results:

# Check ranges for each subset
lapply(h_list, function(x) range(x$height))

Content of the question来源于stack exchange,提问作者stats_noob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 22:37:53