使用Maximal Information Coefficient分析波士顿房价数据集遇minepy维度错误
Hey there, let's sort out this error you're hitting with minepy's MINE!
Why the Error Happens
The compute_score() method from minepy.MINE expects two 1-dimensional arrays (one feature, one target) as inputs. But in your code, you're passing df[col_name[0:13]]—which is a 13-column DataFrame (a 2D structure)—as the first argument. That's exactly why you're getting the "Buffer has wrong number of dimensions" message.
Corrected Code
Instead of passing all features at once, we need to loop through each feature individually and calculate its MIC with the target MEDV. Here's the fixed version:
import numpy as np import pandas as pd from minepy import MINE # 读取数据集 df = pd.read_csv('housing.data', delim_whitespace=True, header=None) col_name = ['CRIM', 'ZN' , 'INDUS', 'CHAS', 'NOX', 'RM', 'AGE', 'DIS', 'RAD', 'TAX', 'PTRATIO', 'B', 'LSTAT', 'MEDV'] df.columns = col_name # 初始化MINE对象 m = MINE() # 循环计算每个特征与MEDV的MIC mic_results = {} for feature in col_name[:-1]: # 排除最后一个目标列MEDV m.compute_score(df[feature], df['MEDV']) mic_results[feature] = m.mic() # 把结果转换成DataFrame,方便查看 mic_df = pd.DataFrame.from_dict(mic_results, orient='index', columns=['MIC']) print(mic_df.sort_values(by='MIC', ascending=False))
What This Does
- We iterate over each feature (excluding
MEDV) and pass each as a 1D Series tocompute_score()alongside the targetMEDV. - We We store each MIC value in a dictionary, then convert it to a sorted DataFrame so you can easily see which features have the highest mutual information with housing prices.
Bonus Tip
If you want to avoid reinitializing the MINE object each time (though it's not strictly necessary), you can just reuse the same m instance as we did here—it works perfectly fine for multiple calculations.
内容的提问来源于stack exchange,提问作者Anx8

