You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中找出数据框中存在超出指定范围数据的列

Solution to Identify Columns with Values Outside Specified Ranges

Let's break down how to fix this problem step by step—your initial approach was on the right track, but it missed two key points: handling continuous ranges instead of discrete values, and mapping each column in df to its specific range in ranges.

First, Let's Restate the Problem Clearly

We have:

  • A data frame df with columns x1, x2, x3
  • A range data frame ranges where each column (y1, y2, y3) corresponds to the min/max allowed values for x1, x2, x3 respectively
  • We need to find columns in df that contain any value outside their corresponding range (expected output: "x1", "x2")

Why Your Initial Code Failed

Your code used %in% ranges which:

  1. Treats ranges as discrete integer sets (e.g., 15:300 only includes integers, missing decimals if they existed)
  2. Uses a single combined range instead of mapping each column to its own unique range
  3. Checks for values inside the range instead of values outside it

Correct Solution (Base R)

First, make sure we map each column in df to its matching range in ranges, then check for out-of-bounds values:

# Define your data
df <- data.frame(x1=c(1,5,9),x2=c(10,20,30),x3=c(20,100,1000))
ranges <- data.frame(y1=c(3,8),y2=c(10,20), y3=c(15,1250))

# Align column names between df and ranges to avoid mismatches
colnames(ranges) <- colnames(df)

# Extract min/max for each column's range
column_ranges <- apply(ranges, 2, function(col) c(min = col[1], max = col[2]))

# Check each column for values outside its range
out_of_range_cols <- colnames(df)[sapply(seq_along(df), function(i) {
  current_col <- df[[i]]
  min_val <- column_ranges["min", i]
  max_val <- column_ranges["max", i]
  # Return TRUE if any value is < min OR > max
  any(current_col < min_val | current_col > max_val)
})]

# View the result
out_of_range_cols
# Output: [1] "x1" "x2"

Alternative Solution (Tidyverse/Purrr)

If you prefer using the tidyverse ecosystem, this approach is more readable and pipe-friendly:

library(purrr)
library(dplyr)

# Define data
df <- data.frame(x1=c(1,5,9),x2=c(10,20,30),x3=c(20,100,1000))
ranges <- data.frame(y1=c(3,8),y2=c(10,20), y3=c(15,1250))

# Align column names
colnames(ranges) <- colnames(df)

# Iterate over each column and check for out-of-range values
out_of_range_cols <- df %>%
  imap(function(col_data, col_name) {
    col_min <- min(ranges[[col_name]])
    col_max <- max(ranges[[col_name]])
    if (any(col_data < col_min | col_data > col_max)) col_name else NA
  }) %>%
  na.omit() %>%
  as.character()

out_of_range_cols
# Output: [1] "x1" "x2"

Key Notes

  • Always align column names (or positions) between df and ranges to ensure you're checking the correct range for each column
  • Use any(value < min | value > max) to check for out-of-bounds values—this works for continuous ranges, not just discrete integers
  • Both methods will handle numeric values (integers or decimals) correctly

内容的提问来源于stack exchange,提问作者Juan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:30:51