解决NumPy元组排序标量转换错误及特征向量载荷输出问题
Hey there, let's tackle your two issues one by one:
一、Fixing the TypeError When Sorting Eigenvalue-Eigenvector Pairs
The error TypeError: only length-1 arrays can be converted to Python scalars pops up because of how np.linalg.eig returns eigenvectors: eig_vecs is a matrix where each column corresponds to an eigenvector, not each row. When you use zip(eig_vals, eig_vecs), you're pairing each scalar eigenvalue with a row from eig_vecs (not the matching eigenvector column), creating tuples of (scalar, 1D array). Converting this list to a numpy array fails because numpy can't treat those 1D arrays as scalars to build a regular array.
Here's the corrected code to properly pair and sort your eigenvalues and eigenvectors:
# Generate list of (eigenvalue, eigenvector) tuples correctly # Transpose eig_vecs so each row is an eigenvector matching the index of eig_vals eig_pairs = list(zip(eig_vals, eig_vecs.T)) # Sort the pairs in descending order of eigenvalues eig_pairs.sort(key=lambda x: x[0], reverse=True) # Optional: Convert back to numpy array if needed (use dtype=object to handle mixed types) eig_pairs = np.array(eig_pairs, dtype=object)
After this, your loop to print sorted eigenvalues will work as expected.
二、Exporting Loadings (or Scores) to a File
First, let's clarify two common terms to ensure we're aligned:
- Variable Loadings: The weight of each original feature on each principal component (calculated as eigenvectors multiplied by the square root of their corresponding eigenvalues).
- Sample Scores: The projection of each data point onto the principal components (calculated by multiplying standardized data with the eigenvector matrix).
1. Export Variable Loadings
# Calculate variable loadings: eigenvectors * sqrt(eigenvalues) loadings = eig_vecs * np.sqrt(eig_vals) # Convert to DataFrame for readability (use original feature names as index) loadings_df = pd.DataFrame( loadings, columns=[f'PC{i+1}' for i in range(len(eig_vals))], index=df_to_save.columns ) # Save to CSV file loadings_df.to_csv('variable_loadings.csv', index=True)
2. Export Sample Principal Component Scores
# Calculate sample scores: standardized data @ eigenvector matrix scores = X_std @ eig_vecs # Convert to DataFrame (retain original index if needed) scores_df = pd.DataFrame( scores, columns=[f'PC{i+1}' for i in range(len(eig_vals))] ) scores_df.index = df_to_save.index # Match original data's index # Save to CSV file scores_df.to_csv('sample_scores.csv', index=True)
If you want to use the sorted eigenvalues/eigenvectors (from the first part of the solution) for loadings/scores, extract them from the sorted eig_pairs first:
# Extract sorted eigenvalues and eigenvectors sorted_eig_vals = np.array([pair[0] for pair in eig_pairs]) sorted_eig_vecs = np.array([pair[1] for pair in eig_pairs]).T # Transpose back to column matrix # Calculate sorted loadings and scores sorted_loadings = sorted_eig_vecs * np.sqrt(sorted_eig_vals) sorted_scores = X_std @ sorted_eig_vecs
内容的提问来源于stack exchange,提问作者James Danowski

