分组处理DataFrame时cut2生成的十分位数区间格式错误求助
group_by() It looks like the issue stems from how you're extracting and processing the interval labels generated by cut2(). Let's break down what's going wrong and fix it step by step.
What's Causing the Problem?
Your current code uses str_extract_all(Value.fc, "\\d+") to pull numbers from the interval strings. The regex \\d+ only matches sequences of digits—so when cut2() generates intervals like (135.0,150], this regex extracts 135, 0, and 150 separately. When you paste those together, you get the malformed 135-0-150 instead of the intended 135-150.
Solution 1: Fix the Regex to Match Full Numeric Values
Instead of extracting just digits, use a regex that matches full floating-point numbers (including optional decimals). Then we can take the first and last extracted values to form our interval:
library(dplyr) library(Hmisc) library(stringr) df <- df %>% group_by(short_desc) %>% mutate( # Generate deciles with cut2 Value.fc = cut2(Value, g=10), # Extract full numeric values (including decimals) Value.fc = factor(sapply(str_extract_all(Value.fc, "\\d+\\.?\\d*"), function(x) { # Take first and last value to form the interval paste(x[1], x[length(x)], collapse="-") })), # Reorder the factor by original Value Value.fc = reorder(Value.fc, Value) ) %>% ungroup()
This regex \\d+\\.?\\d* matches:
\\d+: One or more digits (the integer part)\\.?: An optional decimal point\\d*: Zero or more digits (the decimal part)
If you want to clean up trailing .0 for whole numbers to make labels cleaner, add a str_remove step inside the function:
paste(x[1], x[length(x)], collapse="-") %>% str_remove("\\.0$")
Solution 2: Generate Custom Labels Directly with cut2()
Another cleaner approach is to avoid post-processing entirely by generating custom labels when calling cut2(). We can precompute decile breaks for each group, then create labels manually:
df <- df %>% group_by(short_desc) %>% mutate( # Calculate decile breaks, removing duplicates for tied values breaks = quantile(Value, probs = seq(0, 1, 0.1), na.rm = TRUE) %>% unique(), # Generate intervals with clean custom labels Value.fc = cut(Value, breaks = breaks, include.lowest = TRUE, labels = paste(head(breaks, -1), tail(breaks, -1), sep="-")), # Reorder the factor to match value order Value.fc = reorder(Value.fc, Value) ) %>% ungroup()
This method gives you full control over label formatting without string manipulation. The unique(breaks) step handles cases where some deciles have identical values (common in small or tied datasets).
Testing the Fix
Using your sample data, both solutions will produce the correct Value.fc values:
- For
Value = 137.8(Iowa Corn grain), you'll get135-150instead of135-0-150 - For
Value =13.5(Iowa Corn silage), you'll get13-14.5instead of13-0-14-5
内容的提问来源于stack exchange,提问作者gm007

