You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言data.table:解决cut函数忽略范围外值的土壤湿度统计问题

Hey there! Let's work through this soil humidity (SM) categorization and percentage calculation problem together. I get that dealing with cut() ignoring values outside your breaks is frustrating, but we can fix this with a few parameter tweaks and some tidy data workflows.

Step 1: Understand the cut() Issue

By default, cut() excludes the lowest break value (0 in your case) if you don't specify include.lowest = TRUE, and it treats intervals as left-open/right-closed. For your data, this would mean values exactly equal to 0 get dropped, and we need to make sure the upper limit (0.8) is properly included too.

Step 2: Define Your Breaks and Labels

First, let's formalize your break points and add clear labels for each SM category:

breaks <- seq(0, 0.8, by = 0.2)
# Custom labels to make output readable
category_labels <- c("[0, 0.2]", "(0.2, 0.4]", "(0.4, 0.6]", "(0.6, 0.8]")

Step 3: Fix cut() to Include All Values

Use these parameters in cut() to ensure no valid SM values get excluded:

  • include.lowest = TRUE: Makes the first interval include the minimum break (0)
  • right = TRUE: Keeps intervals right-closed (so values like 0.2, 0.4 fall into the correct left interval)
  • labels = category_labels: Assigns human-readable names to each category

Step 4: Full Workflow with Tidy Data (Using dplyr)

Let's put this all together with a simulated dataset (matching your two-year scenario with NA values) to show the full process:

library(dplyr)
library(lubridate)

# Simulate your SM data (with NA values from sensor failures)
set.seed(123) # For reproducibility
dates <- seq(as.Date("2022-01-01"), as.Date("2023-12-31"), by = "day")
sm_data <- tibble(
  date = dates,
  year = year(date),
  month = month(date, label = TRUE), # Get month names (Jan, Feb, etc.)
  SM = case_when(
    year == 2022 ~ runif(sum(year(date) == 2022), 0, 0.6), # Drier year: 0-0.6
    year == 2023 ~ runif(sum(year(date) == 2023), 0, 0.8)  # Wetter year: 0-0.8
  ) %>% replace(sample(seq_along(.), 50), NA) # Add 50 random NA values
)

# Calculate monthly category percentages
sm_monthly_stats <- sm_data %>%
  filter(!is.na(SM)) %>% # Drop NA values first (since they don't contribute to valid SM readings)
  group_by(year, month) %>%
  mutate(
    sm_category = cut(
      SM,
      breaks = breaks,
      include.lowest = TRUE,
      right = TRUE,
      labels = category_labels
    )
  ) %>%
  count(sm_category, name = "reading_count") %>%
  mutate(
    total_monthly_readings = sum(reading_count),
    percentage = (reading_count / total_monthly_readings) * 100
  ) %>%
  ungroup()

# View the result
head(sm_monthly_stats)

Key Notes:

  • Handling NA Values: We filter out NA values upfront with filter(!is.na(SM)) so they don't skew our percentage calculations. If you want to track NA counts separately, you can skip this step and count() will include NA as a category.
  • Reproducibility: Using set.seed() ensures your simulated data matches the example, but you can remove this when working with your real dataset.
  • Flexibility: If you prefer left-closed intervals instead, just set right = FALSE and adjust your labels accordingly (e.g., "(0, 0.2]", "(0.2, 0.4]", etc.).

This workflow will give you the monthly percentage of SM readings in each of your defined categories, with no valid values excluded by cut().

内容的提问来源于stack exchange,提问作者thiagoveloso

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:38:26