You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DataFrame中更新行名与列名匹配的记录?

Problem Description

I have a large DataFrame where I need to update or fill missing values (marked as XXXX in the example) based on the correspondence between row names and column names. Here's an example:

df <- data.frame(
  ID = c("x", "y", "z"), 
  x = c("1", "0.45", "0.47"),
  y = c("0.45", "1", "0.65"),
  z = c("XXXX", "XXXX", "1")
)

Which produces this DataFrame:

ID    x    y    z
1  x    1 0.45 XXXX
2  y 0.45    1 XXXX
3  z 0.47 0.65    1

The correct values for the XXXX entries should be 0.47 and 0.65, respectively. This is because the value in column x, row z is 0.47, and the value in column y, row z is 0.65. The desired result is a symmetric DataFrame where each element matches the corresponding value at the mirrored row-column position:

ID    x    y    z
1  x    1 0.45 0.47
2  y 0.45    1 0.65
3  z 0.47 0.65    1

I've referenced several Stack Overflow posts but haven't been able to work out a solution.


Solution

Since your DataFrame represents a symmetric matrix (with the ID column acting as row identifiers), you can leverage matrix symmetry to fill the missing values efficiently:

  1. Convert the DataFrame to a numeric matrix (excluding the ID column)
  2. Use the transposed matrix to fill missing entries
  3. Reconstruct the DataFrame with the filled values

Here's the code:

# Load the example DataFrame
df <- data.frame(
  ID = c("x", "y", "z"), 
  x = c("1", "0.45", "0.47"),
  y = c("0.45", "1", "0.65"),
  z = c("XXXX", "XXXX", "1"),
  stringsAsFactors = FALSE
)

# Convert columns to numeric, replacing "XXXX" with NA
df[, -1] <- lapply(df[, -1], function(col) {
  as.numeric(ifelse(col == "XXXX", NA, col))
})

# Extract numeric matrix and set row names to match ID column
mat <- as.matrix(df[, -1])
rownames(mat) <- df$ID

# Fill missing values using symmetric position from transposed matrix
mat[is.na(mat)] <- t(mat)[is.na(mat)]

# Reconstruct the filled DataFrame
filled_df <- cbind(ID = df$ID, as.data.frame(mat))

# Print the result
print(filled_df)

This will output:

ID    x    y    z
1  x 1.00 0.45 0.47
2  y 0.45 1.00 0.65
3  z 0.47 0.65 1.00

Explanation

  • First, we convert character columns to numeric, replacing "XXXX" with NA (R's standard missing value marker).
  • By converting to a matrix and setting row names to match the ID values, we can use the transpose (t(mat)) to access the mirrored position of each missing entry.
  • The line mat[is.na(mat)] <- t(mat)[is.na(mat)] directly replaces every missing value with the corresponding value from the symmetric row-column pair, making use of your data's inherent symmetry.

内容的提问来源于stack exchange,提问作者AOE_player

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 02:57:32