如何在R中利用聚类表为data.frame的百分比值分配对应字母
Got it, let's work through this. You want to take those percentage columns in your databasis data frame and replace each value with a corresponding letter from your cluster table—like assigning 'C' to Aa's 0.0833 in X200202. Here's a straightforward, flexible way to do this:
Step 1: Define Your Original Data Frame
First, let's make sure we have your data frame set up properly (I filled in an extra column for example):
# Original databasis data frame company <- c("Aa", "Bb", "Cc", "Dd") sector <- c("B", "A", "E", "Z") lastmarketcap <- c(18, 4571, 122, 239) X200202 <- c(0.0833, 0.0002, 0.00, 0.0085) X200203 <- c(0.021, 0.067, 0.0005, 0.09) # Example additional date column databasis <- data.frame(company, sector, lastmarketcap, X200202, X200203)
Step 2: Set Up Your Cluster Mapping Table
You'll need to define your cluster rules clearly. Since your example assigns 'C' to 0.0833, I'll use a range-based mapping (adjust this to match your actual cluster table):
# Cluster mapping table (customize thresholds and letters to your needs) cluster_table <- data.frame( lower_threshold = c(-Inf, 0.001, 0.05), upper_threshold = c(0.001, 0.05, Inf), assigned_letter = c("A", "B", "C") )
This means:
- Values < 0.001 get 'A'
- Values between 0.001 and 0.05 get 'B'
- Values >=0.05 get 'C'
Step 3: Create a Mapping Function
This helper function will take a percentage value and return the correct letter from your cluster table:
map_to_letter <- function(percent_value) { # Find which cluster the value falls into cluster_row <- which(percent_value >= cluster_table$lower_threshold & percent_value < cluster_table$upper_threshold) # Return the corresponding letter cluster_table$assigned_letter[cluster_row] }
Step 4: Apply the Mapping to Your Data Frame
Using dplyr (a common R package for data manipulation), we can apply this function to all your date-based percentage columns (the ones starting with 'X'):
# Load dplyr if you haven't already library(dplyr) # Create the new data frame with mapped letters new_databasis <- databasis %>% # Apply the mapping to all columns starting with 'X' mutate(across(starts_with("X"), ~sapply(.x, map_to_letter))) %>% # Optional: Rename the mapped columns to add a "_letter" suffix for clarity rename_with(~paste0(.x, "_letter"), starts_with("X"))
Step 5: Check the Result
If you print new_databasis, you'll see:
- Aa's X200202_letter is 'C' (matches your example)
- Bb's X200202_letter is 'A'
- Cc's X200202_letter is 'A'
- Dd's X200202_letter is 'B'
Alternative: Exact Value Mapping
If your cluster table uses exact percentage values instead of ranges, use a lookup vector instead:
# Exact value to letter mapping (adjust to your actual values) cluster_lookup <- c("0.0833" = "C", "0.0002" = "A", "0.00" = "A", "0.0085" = "B") new_databasis <- databasis %>% mutate(across(starts_with("X"), ~recode(as.character(.x), !!!cluster_lookup)))
Just adjust the cluster rules to fit your specific table, and this should work smoothly!
内容的提问来源于stack exchange,提问作者Robin_Hcp

