如何用Base R拆分列、关联数据并保留高频对应值?
Let's walk through how to solve this step by step—your initial code only handled the Paths column, so we'll fix that by linking each split path directly to its corresponding color.
Step 1: Recreate the Original Data Frame (for reproducibility)
First, let's make sure we're working with the exact data you provided:
df <- data.frame( Names = c("IL-1", "IL-6", "TNF", "CCL2", "IL-1", "PAI1", "IL-12", "IL-8", "CTGF", "TGF"), Symbol = c("CP1", "CFT5", "DFR4", "FGTZ", "SED", "SWA", "SSS", "SDE", "SDFR", "FDGT"), GeneID = c(3553, 3569, 7124, 6347, 3552, 5054, 3593, 3576, 1490, 7046), Description = rep("Receptor", 10), Paths = c("Path1|Path2|Path5", "Path3|Path1|Path2", "Path4|Path3|Path1", "Path4|Path5|Path2", "Path6|Path5|Path3", "Path1", "Path1|Path2", "Path1|Path3|Path5", "Path4|Path5|Path1", "Path5|Path3"), Colors = c("Green|Blue|Pink", "Red|Green|Blue", "Yellow|Red|Green", "Yellow|Pink|Blue", "Purple|Pink|Red", "Green", "Green|Blue", "Green|Red|Pink", "Yellow|Pink|Green", "Pink|Red"), stringsAsFactors = FALSE )
Step 2: Split and Pair Paths with Their Corresponding Colors
The critical part here is ensuring each split Path stays linked to its matching Color from the same row. We'll use mapply to process each row's Paths and Colors together, then combine all results into a single long data frame:
# Split each row's Paths and Colors, then pair them into a data frame path_color_pairs <- mapply(function(p, c) { data.frame( Path = strsplit(p, "|", fixed = TRUE)[[1]], Color = strsplit(c, "|", fixed = TRUE)[[1]], stringsAsFactors = FALSE ) }, df$Paths, df$Colors, SIMPLIFY = FALSE) # Combine all paired data frames into one unified table path_color_df <- do.call(rbind, path_color_pairs)
Step 3: Count Path Frequencies and Find Top Colors
Next, we'll count how often each path appears, then determine the most frequent color associated with each path:
# Count path occurrences and sort in descending order path_counts <- table(path_color_df$Path) path_counts_df <- data.frame( Path = names(path_counts), Count = as.numeric(path_counts), stringsAsFactors = FALSE ) path_counts_df <- path_counts_df[order(-path_counts_df$Count), ] # Function to get the most frequent color for a given path get_top_color <- function(path) { color_table <- table(path_color_df$Color[path_color_df$Path == path]) names(color_table)[which.max(color_table)] # Pick the highest-frequency color } # Add the top color to each path entry path_counts_df$Color <- sapply(path_counts_df$Path, get_top_color)
Step 4: Finalize the Output Data Frame
Finally, we'll adjust the column names and structure to match your desired df2:
# Rename columns and keep only the required fields df2 <- path_counts_df[, c("Path", "Color")] colnames(df2) <- c("Paths", "Colors") # View the final result print(df2)
This will produce exactly the output you wanted:
Paths Colors Path1 Path1 Green Path5 Path5 Pink Path3 Path3 Red Path2 Path2 Blue Path4 Path4 Yellow Path6 Path6 Purple
内容的提问来源于stack exchange,提问作者Biocrazy

