You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中基于另一data.frame的text.values创建相关性data.frame?

Hey there! Let's walk through how to build that correlation data frame step by step—no need to feel stuck anymore. First, let's align on your data structure and what "common changes" means, then dive into actionable code.

Step 1: Define Your Data Structure

First, let's assume your data frame looks something like this (I'll make a sample with 4 periods and 3 companies to match your example):

# Set a seed for reproducibility
set.seed(123)
df <- data.frame(
  Period = 1:4,
  Aa = sample(c("A", "B", "C", "D", "E", "F"), 4, replace = TRUE),
  Cc = sample(c("A", "B", "C", "D", "E", "F"), 4, replace = TRUE),
  Dd = sample(c("A", "B", "C", "D", "E", "F"), 4, replace = TRUE)
)

This gives us a table where each row is a period, and each column holds a company's value (A-F).

Step 2: Clarify "Common Changes"

From your example, it sounds like you mean number of periods where two companies have the exact same value divided by total periods. If that's not your actual definition (e.g., you mean they changed in the same direction between periods), I’ll adjust the code later—just pick the version that fits your needs!

Step 3: Calculate Correlations for All Company Pairs

We’ll use base R functions to handle this without extra packages:

First, isolate company columns

# Exclude the Period column to get just company names
companies <- colnames(df)[-1]

Generate all unique company pairs

# Create every possible pair of companies (no duplicates)
company_pairs <- combn(companies, 2, simplify = FALSE)

Compute correlation for each pair

# Loop through each pair to calculate common changes and correlation
corr_results <- lapply(company_pairs, function(pair) {
  # Count how many periods the two companies have matching values
  common_count <- sum(df[[pair[1]]] == df[[pair[2]]])
  # Divide by total periods to get the correlation ratio
  corr_value <- common_count / nrow(df)
  
  # Return a small data frame with the pair and their correlation
  data.frame(
    Company1 = pair[1],
    Company2 = pair[2],
    Correlation = round(corr_value, 2)
  )
})

# Combine all results into one clean data frame
corr_df <- do.call(rbind, corr_results)

Step 4: Convert to a Symmetric Correlation Matrix (Optional)

If you want a more readable matrix where rows and columns are companies (like a traditional correlation table), reshape the results like this:

# Create an empty matrix with company names as rows/columns
corr_matrix <- matrix(NA, nrow = length(companies), ncol = length(companies),
                      dimnames = list(companies, companies))

# A company's correlation with itself is always 1
diag(corr_matrix) <- 1

# Fill in the matrix with our calculated correlations
for (i in 1:nrow(corr_df)) {
  corr_matrix[corr_df$Company1[i], corr_df$Company2[i]] <- corr_df$Correlation[i]
  corr_matrix[corr_df$Company2[i], corr_df$Company1[i]] <- corr_df$Correlation[i]
}

# Convert back to a data frame if needed
corr_matrix_df <- as.data.frame(corr_matrix)

If "Common Changes" Means Direction of Movement

If you actually mean that two companies changed in the same direction between periods (e.g., Aa went from A→B and Cc went from C→D, both increasing), use this adjusted code:

# Convert A-F to numeric values (1-6) to measure changes
df_numeric <- lapply(companies, function(comp) {
  match(df[[comp]], LETTERS[1:6])
})
names(df_numeric) <- companies

# Calculate the change between consecutive periods for each company
company_changes <- lapply(df_numeric, function(vals) {
  diff(vals) # Gives a vector of length 3 for 4 periods
})

# Recalculate correlations based on matching change directions
corr_results_direction <- lapply(company_pairs, function(pair) {
  # Count how many times the change direction (positive/negative/zero) matches
  common_changes <- sum(sign(company_changes[[pair[1]]]) == sign(company_changes[[pair[2]]]))
  # Divide by number of change periods (total periods - 1)
  corr_value <- common_changes / (nrow(df) - 1)
  
  data.frame(
    Company1 = pair[1],
    Company2 = pair[2],
    Correlation = round(corr_value, 2)
  )
})

corr_df_direction <- do.call(rbind, corr_results_direction)

Just swap in whichever definition fits your actual use case!

内容的提问来源于stack exchange,提问作者Robin_Hcp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:29:10