如何使用Pandas命令将df_score_2转换为df_score_1的全矩阵格式?
Got it, let's walk through how to turn your long-form df_score_2 into the square matrix structure of df_score_1. Here's a step-by-step breakdown with Pandas commands:
Step 1: Fix Column Names (if needed)
First, make sure your df_score_2 has clear column names. If you loaded it from a text file without headers, assign meaningful names first:
import pandas as pd # Assuming your df_score_2 looks like this after loading from the text file df_score_2 = pd.DataFrame([ ['A', 'B', 1], ['A', 'C', 1], ['A', 'D', 2], ['B', 'C', 5], ['B', 'D', 1] ]) # Assign descriptive column names df_score_2.columns = ['row', 'col', 'score']
Step 2: Add Reverse Entries (for Symmetric Matrix)
Since df_score_1 is a symmetric matrix (e.g., A→B = 1 and B→A = 1), we need to add the reverse pairs that are missing from df_score_2:
# Create reverse entries by swapping the row and col columns reverse_entries = df_score_2.rename(columns={'row': 'col', 'col': 'row'}) # Combine original and reverse entries into one DataFrame full_entries = pd.concat([df_score_2, reverse_entries], ignore_index=True)
Step 3: Pivot to Square Matrix
Use Pandas' pivot method to reshape the long-form data into a wide matrix:
# Pivot the data: row becomes index, col becomes columns, score fills the cells matrix_df = full_entries.pivot(index='row', columns='col', values='score')
Step 4: Fill Missing Values and Set Diagonal to 0
Fill any empty cells (where no score exists) with 0, and ensure the diagonal (same row/column) is 0 to match df_score_1:
# Replace NaN values with 0 matrix_df = matrix_df.fillna(0) # Manually set diagonal elements to 0 (in case any slipped through) for idx in matrix_df.index: matrix_df.loc[idx, idx] = 0
Step 5: Reindex to Match df_score_1's Order
Finally, reorder the rows and columns to exactly match the sequence in df_score_1 (A, B, C, D):
matrix_df = matrix_df.reindex(index=['A', 'B', 'C', 'D'], columns=['A', 'B', 'C', 'D'])
Final Result
Your matrix_df will now have the exact structure of df_score_1:
A B C D A 0 1 1 2 B 1 0 5 1 C 1 5 0 0 D 2 1 0 0
(Note: The values differ from your sample df_score_1 because df_score_2 uses different non-zero scores, but the matrix structure is identical.)
内容的提问来源于stack exchange,提问作者d..b

