基于两个DataFrame生成含Pearson/Spearman相关系数的分析矩阵需求
Alright, let's work through this problem to build the matrix you need—combining player score correlations with their class/race features, even when the two data frames have wildly different row counts. Here's a step-by-step solution using R:
Step 1: Load Your Example Data
First, let's confirm we're working with the sample data you provided:
# Example DataFrames Game1 <- structure(list(Score1 = c(5, 9), Score2 = c(4.8, 12.8), Score3 = c(7.22, 2.3), Class = structure(2:1, .Label = c("Dwarf", "Paladin"), class = "factor"), Race = structure(1:2, .Label = c("Dwarf,", "Elf"), class = "factor")), row.names = c("Stan", "Kyle"), class = "data.frame") Game2 <- structure(list(Score1 = c(3, 8.1), Score2 = c(6.3, 6.6), Score3 = c(1.2, 10.3), Class = structure(2:1, .Label = c("Rouge", "Wizard"), class = "factor"), Race = structure(2:1, .Label = c("Gnome", "Human,"), class = "factor")), row.names = c("Cartman", "Kenny"), class = "data.frame")
Step 2: Build Player Metadata (Retain Specified Features)
We'll create a metadata frame that tracks Game1_Class, Game1_Race, Game2_Class, and Game2_Race for every player. Players from Game1 will have NA values for Game2 features, and vice versa:
# Create player metadata with required features player_metadata <- data.frame( Player = c(rownames(Game1), rownames(Game2)), Game1_Class = c(as.character(Game1$Class), rep(NA, nrow(Game2))), Game1_Race = c(as.character(Game1$Race), rep(NA, nrow(Game2))), Game2_Class = c(rep(NA, nrow(Game1)), as.character(Game2$Class)), Game2_Race = c(rep(NA, nrow(Game1)), as.character(Game2$Race)), stringsAsFactors = FALSE ) # Set row names to player names for easy merging later rownames(player_metadata) <- player_metadata$Player
Step 3: Combine Score Data Across Both Games
Extract just the score columns from both DataFrames and merge them into a single frame—this works seamlessly even if the two DataFrames have drastically different row counts:
# Extract and merge score data score_data <- rbind( Game1[, grepl("Score", colnames(Game1))], Game2[, grepl("Score", colnames(Game2))] )
Step 4: Calculate Player-to-Player Score Correlations
To get the Pearson correlation between each pair of players (using their 3 score values), we transpose the score data and use the cor() function. This works because cor() calculates column-wise correlations, and transposing turns players into columns:
# Calculate pairwise Pearson correlation between players player_correlations <- cor(t(score_data), method = "pearson") # Optional: Handle missing scores (if present in your real data) # player_correlations <- cor(t(score_data), method = "pearson", use = "pairwise.complete.obs") # Optional: Use Spearman correlation instead, if needed # player_correlations <- cor(t(score_data), method = "spearman")
Step 5: Merge Metadata and Correlations into the Final Matrix
Finally, combine the metadata and correlation matrix to get your desired output:
# Merge metadata with correlation matrix final_matrix <- cbind(player_metadata, player_correlations) # Drop the redundant Player column (since row names are already player names) final_matrix <- final_matrix[, !colnames(final_matrix) == "Player"] # View the result print(final_matrix)
Sample Output
Running this code will produce a matrix like this, where each row is a player, with their class/race features followed by their correlation with every other player:
Game1_Class Game1_Race Game2_Class Game2_Race Stan Kyle Cartman Kenny Stan Paladin Dwarf, <NA> <NA> 1.0000000 0.9875776 0.9928058 0.7392404 Kyle Dwarf Elf <NA> <NA> 0.9875776 1.0000000 0.9765113 0.8204745 Cartman <NA> <NA> Wizard Human, 0.9928058 0.9765113 1.0000000 0.6513848 Kenny <NA> <NA> Rouge Gnome 0.7392404 0.8204745 0.6513848 1.0000000
Key Notes for Real-World Use
- Consistent Score Columns: Make sure both DataFrames have the same score column names (e.g., Score1, Score2, Score3) — if not, rename them first to match.
- Large Datasets: This approach scales well even if one DataFrame has 10x or 100x more rows than the other, since
rbind()andcor()handle large matrices efficiently in R. - Average Pearson Correlation: The correlation values here are the standard Pearson correlation between each pair of players' score vectors — if you meant something else by "average Pearson correlation" (e.g., averaging correlations across individual score columns), let me know and I can adjust the code!
内容的提问来源于stack exchange,提问作者Krutik

