机器学习模型得分字典的各预处理类别平均值计算及TypeError报错解决
问题分析与解决方案
嘿,我一眼就看出你遇到的TypeError问题根源啦——你的字典里所有模型得分都是字符串类型,而mean()函数只能处理数值(比如int或float),没法直接把字符串当作数字来计算平均值,这就是报错的核心原因。
修正思路
核心解决步骤就是:把每个字符串格式的得分转换成浮点型(float),然后再计算均值。我帮你优化了代码,不仅修复了错误,还加上了筛选最高平均得分预处理方法的逻辑,正好满足你的需求。
完整修正代码
from statistics import mean # 你的原始得分字典 model_scores_for_datasets = { "Unprocessed": {"Logistic Regression": "0.967", "Support Vector Machine": "0.967", "Decision Tree": "0.933", "Random Forest": "0.933", "LinearDiscriminant": "1.000", "K-Nearest Neighbour": "1.000", "Naive Bayes": "0.967", "XGBoost": "0.933"}, "Standardisation": {"Logistic Regression": "0.933", "Support Vector Machine": "0.967", "Decision Tree": "0.933", "Random Forest": "0.967", "LinearDiscriminant": "0.967", "K-Nearest Neighbour": "0.967", "Naive Bayes": "0.967", "XGBoost": "0.933"}, "Normalisation": {"Logistic Regression": "0.967", "Support Vector Machine": "0.967", "Decision Tree": "0.933", "Random Forest": "0.967", "LinearDiscriminant": "0.967", "K-Nearest Neighbour": "0.967", "Naive Bayes": "0.967", "XGBoost": "0.933"}, "Rescale": {"Logistic Regression": "0.967", "Support Vector Machine": "0.967", "Decision Tree": "0.933", "Random Forest": "0.933", "LinearDiscriminant": "0.967", "K-Nearest Neighbour": "0.967", "Naive Bayes": "0.967", "XGBoost": "0.933"} } # 存储每个预处理类别的平均得分 preprocessing_avg_scores = {} # 遍历计算每个类别的均值 for preprocessing_type, model_scores in model_scores_for_datasets.items(): # 将字符串得分转换为浮点型数值 numeric_scores = [float(score) for score in model_scores.values()] # 计算均值并保留三位小数 avg_score = mean(numeric_scores) preprocessing_avg_scores[preprocessing_type] = round(avg_score, 3) print(f"Average for {preprocessing_type} is {avg_score:.3f}") # 筛选出平均得分最高的预处理方法 best_preprocessing = max(preprocessing_avg_scores.items(), key=lambda x: x[1]) print(f"\n✅ The best preprocessing method is '{best_preprocessing[0]}' with an average score of {best_preprocessing[1]:.3f}")
关键细节说明
- 类型转换:用列表推导式
[float(score) for score in model_scores.values()]批量把字符串转成浮点型,这是解决报错的核心步骤 - 均值计算:使用Python标准库
statistics里的mean()函数,如果你习惯用NumPy,也可以替换成np.mean(numeric_scores)(记得先导入numpy) - 最优方法筛选:通过
max()函数结合lambda表达式,直接从平均得分字典里找出得分最高的预处理类别
内容的提问来源于stack exchange,提问作者Francesca C
相关产品推荐
相关产品推荐

