You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

分组处理DataFrame时cut2生成的十分位数区间格式错误求助

Fixing Malformed Decile Intervals After Grouping with group_by()

It looks like the issue stems from how you're extracting and processing the interval labels generated by cut2(). Let's break down what's going wrong and fix it step by step.

What's Causing the Problem?

Your current code uses str_extract_all(Value.fc, "\\d+") to pull numbers from the interval strings. The regex \\d+ only matches sequences of digits—so when cut2() generates intervals like (135.0,150], this regex extracts 135, 0, and 150 separately. When you paste those together, you get the malformed 135-0-150 instead of the intended 135-150.

Solution 1: Fix the Regex to Match Full Numeric Values

Instead of extracting just digits, use a regex that matches full floating-point numbers (including optional decimals). Then we can take the first and last extracted values to form our interval:

library(dplyr)
library(Hmisc)
library(stringr)

df <- df %>%
  group_by(short_desc) %>%
  mutate(
    # Generate deciles with cut2
    Value.fc = cut2(Value, g=10),
    # Extract full numeric values (including decimals)
    Value.fc = factor(sapply(str_extract_all(Value.fc, "\\d+\\.?\\d*"), function(x) {
      # Take first and last value to form the interval
      paste(x[1], x[length(x)], collapse="-")
    })),
    # Reorder the factor by original Value
    Value.fc = reorder(Value.fc, Value)
  ) %>%
  ungroup()

This regex \\d+\\.?\\d* matches:

  • \\d+: One or more digits (the integer part)
  • \\.?: An optional decimal point
  • \\d*: Zero or more digits (the decimal part)

If you want to clean up trailing .0 for whole numbers to make labels cleaner, add a str_remove step inside the function:

paste(x[1], x[length(x)], collapse="-") %>% str_remove("\\.0$")

Solution 2: Generate Custom Labels Directly with cut2()

Another cleaner approach is to avoid post-processing entirely by generating custom labels when calling cut2(). We can precompute decile breaks for each group, then create labels manually:

df <- df %>%
  group_by(short_desc) %>%
  mutate(
    # Calculate decile breaks, removing duplicates for tied values
    breaks = quantile(Value, probs = seq(0, 1, 0.1), na.rm = TRUE) %>% unique(),
    # Generate intervals with clean custom labels
    Value.fc = cut(Value, breaks = breaks, include.lowest = TRUE,
                   labels = paste(head(breaks, -1), tail(breaks, -1), sep="-")),
    # Reorder the factor to match value order
    Value.fc = reorder(Value.fc, Value)
  ) %>%
  ungroup()

This method gives you full control over label formatting without string manipulation. The unique(breaks) step handles cases where some deciles have identical values (common in small or tied datasets).

Testing the Fix

Using your sample data, both solutions will produce the correct Value.fc values:

  • For Value = 137.8 (Iowa Corn grain), you'll get 135-150 instead of 135-0-150
  • For Value =13.5 (Iowa Corn silage), you'll get 13-14.5 instead of 13-0-14-5

内容的提问来源于stack exchange,提问作者gm007

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:52:10