R语言处理iris数据集:生成分箱列与计数列并按物种汇总
Got it, let's work through this together to get the full result for all three iris species. You want to group by Species, bin Sepal.Length into 4 intervals (labeled D1-D4), and add corresponding count columns (C1-C4) for each bin per species. Here's a complete, reproducible solution using R's tidyverse tools:
Step 1: Load Required Packages
First, we'll use dplyr for grouping and data manipulation, and tidyr to reshape our data into the wide format you need:
library(dplyr) library(tidyr)
Step 2: Full Data Processing Code
This code will handle all species at once, generate the bins, count samples per bin, and format the output to match your requested structure:
# Process iris data to create bin intervals and counts per species iris_binned <- iris %>% # Group data by each iris species group_by(Species) %>% # Create 4 equal-width bins for Sepal.Length, labeled D1-D4 mutate(Sepal_Bin = cut( Sepal.Length, breaks = 4, labels = paste0("D", 1:4), include.lowest = TRUE # Ensures the smallest value is included in the first bin )) %>% # Count how many samples fall into each bin per species count(Sepal_Bin, name = "Sample_Count") %>% # Reshape data to wide format: separate bin intervals and counts into D1-D4/C1-C4 pivot_wider( names_from = Sepal_Bin, values_from = c(Sepal_Bin, Sample_Count), names_glue = "{ifelse(.value == 'Sepal_Bin', '', 'C')}{gsub('D', '', Sepal_Bin)}" ) %>% # Rename columns to match D1-D4 and C1-C4 exactly rename_with(~ gsub("Sepal_Bin", "D", .x), starts_with("Sepal_Bin")) %>% # Replace any NA counts with 0 (for bins with no samples in a species) mutate(across(starts_with("C"), ~ replace_na(.x, 0))) %>% # Reorder columns to put intervals first, then counts select(Species, D1, D2, D3, D4, C1, C2, C3, C4) # View the final result print(iris_binned)
Step 3: Example Output
When you run the code above, you'll get a table like this (bin intervals may vary slightly based on your exact data, but the structure will match):
# A tibble: 3 × 9 Species D1 D2 D3 D4 C1 C2 C3 C4 <fct> <fct> <fct> <fct> <fct> <dbl> <dbl> <dbl> <dbl> 1 setosa [4.3,4.82] (4.82,5.35] (5.35,5.88] (5.88,6.4] 11 23 14 2 2 versicolor [4.9,5.62] (5.62,6.35] (6.35,7.08] (7.08,7.8] 7 23 16 4 3 virginica [4.9,5.92] (5.92,6.95] (6.95,7.98] (7.98,9] 1 12 28 9
Notes on Customization
- If you want quantile-based bins (where each bin has roughly the same number of samples, instead of equal width), replace the
cutline with:Sepal_Bin = cut( Sepal.Length, breaks = quantile(Sepal.Length, probs = 0:4/4), labels = paste0("D", 1:4), include.lowest = TRUE ) - The
include.lowest = TRUEparameter ensures the minimumSepal.Lengthvalue isn't excluded from the first bin.
内容的提问来源于stack exchange,提问作者lolo

