如何在R中基于另一data.frame的text.values创建相关性data.frame?
Hey there! Let's walk through how to build that correlation data frame step by step—no need to feel stuck anymore. First, let's align on your data structure and what "common changes" means, then dive into actionable code.
Step 1: Define Your Data Structure
First, let's assume your data frame looks something like this (I'll make a sample with 4 periods and 3 companies to match your example):
# Set a seed for reproducibility set.seed(123) df <- data.frame( Period = 1:4, Aa = sample(c("A", "B", "C", "D", "E", "F"), 4, replace = TRUE), Cc = sample(c("A", "B", "C", "D", "E", "F"), 4, replace = TRUE), Dd = sample(c("A", "B", "C", "D", "E", "F"), 4, replace = TRUE) )
This gives us a table where each row is a period, and each column holds a company's value (A-F).
Step 2: Clarify "Common Changes"
From your example, it sounds like you mean number of periods where two companies have the exact same value divided by total periods. If that's not your actual definition (e.g., you mean they changed in the same direction between periods), I’ll adjust the code later—just pick the version that fits your needs!
Step 3: Calculate Correlations for All Company Pairs
We’ll use base R functions to handle this without extra packages:
First, isolate company columns
# Exclude the Period column to get just company names companies <- colnames(df)[-1]
Generate all unique company pairs
# Create every possible pair of companies (no duplicates) company_pairs <- combn(companies, 2, simplify = FALSE)
Compute correlation for each pair
# Loop through each pair to calculate common changes and correlation corr_results <- lapply(company_pairs, function(pair) { # Count how many periods the two companies have matching values common_count <- sum(df[[pair[1]]] == df[[pair[2]]]) # Divide by total periods to get the correlation ratio corr_value <- common_count / nrow(df) # Return a small data frame with the pair and their correlation data.frame( Company1 = pair[1], Company2 = pair[2], Correlation = round(corr_value, 2) ) }) # Combine all results into one clean data frame corr_df <- do.call(rbind, corr_results)
Step 4: Convert to a Symmetric Correlation Matrix (Optional)
If you want a more readable matrix where rows and columns are companies (like a traditional correlation table), reshape the results like this:
# Create an empty matrix with company names as rows/columns corr_matrix <- matrix(NA, nrow = length(companies), ncol = length(companies), dimnames = list(companies, companies)) # A company's correlation with itself is always 1 diag(corr_matrix) <- 1 # Fill in the matrix with our calculated correlations for (i in 1:nrow(corr_df)) { corr_matrix[corr_df$Company1[i], corr_df$Company2[i]] <- corr_df$Correlation[i] corr_matrix[corr_df$Company2[i], corr_df$Company1[i]] <- corr_df$Correlation[i] } # Convert back to a data frame if needed corr_matrix_df <- as.data.frame(corr_matrix)
If "Common Changes" Means Direction of Movement
If you actually mean that two companies changed in the same direction between periods (e.g., Aa went from A→B and Cc went from C→D, both increasing), use this adjusted code:
# Convert A-F to numeric values (1-6) to measure changes df_numeric <- lapply(companies, function(comp) { match(df[[comp]], LETTERS[1:6]) }) names(df_numeric) <- companies # Calculate the change between consecutive periods for each company company_changes <- lapply(df_numeric, function(vals) { diff(vals) # Gives a vector of length 3 for 4 periods }) # Recalculate correlations based on matching change directions corr_results_direction <- lapply(company_pairs, function(pair) { # Count how many times the change direction (positive/negative/zero) matches common_changes <- sum(sign(company_changes[[pair[1]]]) == sign(company_changes[[pair[2]]])) # Divide by number of change periods (total periods - 1) corr_value <- common_changes / (nrow(df) - 1) data.frame( Company1 = pair[1], Company2 = pair[2], Correlation = round(corr_value, 2) ) }) corr_df_direction <- do.call(rbind, corr_results_direction)
Just swap in whichever definition fits your actual use case!
内容的提问来源于stack exchange,提问作者Robin_Hcp

