在R语言中找出数据框中存在超出指定范围数据的列
Solution to Identify Columns with Values Outside Specified Ranges
Let's break down how to fix this problem step by step—your initial approach was on the right track, but it missed two key points: handling continuous ranges instead of discrete values, and mapping each column in df to its specific range in ranges.
First, Let's Restate the Problem Clearly
We have:
- A data frame
dfwith columnsx1,x2,x3 - A range data frame
rangeswhere each column (y1,y2,y3) corresponds to the min/max allowed values forx1,x2,x3respectively - We need to find columns in
dfthat contain any value outside their corresponding range (expected output:"x1", "x2")
Why Your Initial Code Failed
Your code used %in% ranges which:
- Treats ranges as discrete integer sets (e.g.,
15:300only includes integers, missing decimals if they existed) - Uses a single combined range instead of mapping each column to its own unique range
- Checks for values inside the range instead of values outside it
Correct Solution (Base R)
First, make sure we map each column in df to its matching range in ranges, then check for out-of-bounds values:
# Define your data df <- data.frame(x1=c(1,5,9),x2=c(10,20,30),x3=c(20,100,1000)) ranges <- data.frame(y1=c(3,8),y2=c(10,20), y3=c(15,1250)) # Align column names between df and ranges to avoid mismatches colnames(ranges) <- colnames(df) # Extract min/max for each column's range column_ranges <- apply(ranges, 2, function(col) c(min = col[1], max = col[2])) # Check each column for values outside its range out_of_range_cols <- colnames(df)[sapply(seq_along(df), function(i) { current_col <- df[[i]] min_val <- column_ranges["min", i] max_val <- column_ranges["max", i] # Return TRUE if any value is < min OR > max any(current_col < min_val | current_col > max_val) })] # View the result out_of_range_cols # Output: [1] "x1" "x2"
Alternative Solution (Tidyverse/Purrr)
If you prefer using the tidyverse ecosystem, this approach is more readable and pipe-friendly:
library(purrr) library(dplyr) # Define data df <- data.frame(x1=c(1,5,9),x2=c(10,20,30),x3=c(20,100,1000)) ranges <- data.frame(y1=c(3,8),y2=c(10,20), y3=c(15,1250)) # Align column names colnames(ranges) <- colnames(df) # Iterate over each column and check for out-of-range values out_of_range_cols <- df %>% imap(function(col_data, col_name) { col_min <- min(ranges[[col_name]]) col_max <- max(ranges[[col_name]]) if (any(col_data < col_min | col_data > col_max)) col_name else NA }) %>% na.omit() %>% as.character() out_of_range_cols # Output: [1] "x1" "x2"
Key Notes
- Always align column names (or positions) between
dfandrangesto ensure you're checking the correct range for each column - Use
any(value < min | value > max)to check for out-of-bounds values—this works for continuous ranges, not just discrete integers - Both methods will handle numeric values (integers or decimals) correctly
内容的提问来源于stack exchange,提问作者Juan
相关产品推荐
相关产品推荐

