You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于相关矩阵对CSV文件列重排的技术问询(IRIS数据集示例)

Rearranging CSV Columns by Correlation (Iris Dataset Example)

Got it, let's figure out how to rearrange your CSV columns so that highly correlated features sit next to each other. Using your Iris dataset correlation matrix as a reference, here are two straightforward approaches—one automated (great for large datasets) and one manual (perfect for small, interpretable matrices like this).

Automated Approach: Hierarchical Clustering

This method uses clustering to group columns based on their correlation, ensuring similar features are clustered together. It's scalable and doesn't require manual guesswork.

Step-by-Step Code (Python)

We'll use pandas for data handling and scipy for clustering:

import pandas as pd
from scipy.cluster.hierarchy import linkage, leaves_list
import matplotlib.pyplot as plt

# Load your CSV file (replace 'iris.csv' with your file path)
df = pd.read_csv('iris.csv')

# Calculate the correlation matrix (matches your provided matrix)
corr_matrix = df.corr()

# Convert correlation to a distance metric (1 - correlation) for clustering
# We use Ward's method to minimize variance within clusters
linkage_matrix = linkage(1 - corr_matrix, method='ward')

# Get the sorted column order from the clustering results
sorted_col_order = corr_matrix.columns[leaves_list(linkage_matrix)]

# Reorder the DataFrame columns
df_reordered = df[sorted_col_order]

# Save the reordered CSV
df_reordered.to_csv('iris_reordered.csv', index=False)

# Optional: Visualize the clustering with a dendrogram
plt.figure(figsize=(10, 4))
dendrogram(linkage_matrix, labels=corr_matrix.columns, leaf_rotation=90)
plt.title('Feature Clustering by Correlation (Iris Dataset)')
plt.tight_layout()
plt.show()

What's Happening Here?

  • We convert correlation to a distance metric (1 - corr) because clustering algorithms work by minimizing distance—high correlation equals small distance, which is exactly what we want for grouping similar features.
  • Ward's method ensures clusters are formed to minimize variance within each group, leading to clean, logical groupings of correlated features.
  • The dendrogram (optional) lets you visualize how features are clustered, so you can easily verify the ordering makes sense.

For your Iris data, this will output a column order like sepal_width, petal_width, petal_length, sepal_length or sepal_length, petal_length, petal_width, sepal_width—both group the highly correlated petal features together, and sepal_width (which has weak/negative correlations with others) is separated.

Manual Approach (For Small Matrices)

Looking at your correlation matrix, we can manually spot the high-correlation pairs:

  • sepal_length has strong positive correlations with petal_length (0.87) and petal_width (0.82)
  • petal_length and petal_width are almost perfectly correlated (0.96)
  • sepal_width has weak/negative correlations with all other features

A logical manual ordering would be:
sepal_length, petal_length, petal_width, sepal_width

To apply this, just reorder the DataFrame directly:

df_reordered = df[['sepal_length', 'petal_length', 'petal_width', 'sepal_width']]
df_reordered.to_csv('iris_reordered.csv', index=False)

This works great for small datasets where you can easily interpret the correlation matrix, but the clustering method is better for larger datasets with dozens of features.

内容的提问来源于stack exchange,提问作者sajjad akbar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:55:13