如何调整sklearn classification_report列间距保持行列对齐?
解决sklearn长类名导致classification_report对齐错乱的方法
sklearn的classification_report默认按半角字符宽度计算列预留长度,遇到中文、日文这类全角长类名时,会因为宽度计算不足、列间距不够出现行列错位,官方确实没有提供直接调整列间距的参数,可以通过以下几种方案解决:
方案1:输出结构化字典自定义格式化(最稳定推荐)
不要直接使用默认生成的字符串报告,传入output_dict=True拿到结构化的指标结果,自行控制每列宽度和间距,可自动适配最长类名,完全避免错位。
示例代码:
from sklearn.metrics import classification_report # 替换为实际的真实标签、预测标签、类别名列表 y_true = [] y_pred = [] class_names = ["カップ", "セット", "パウチ", "ビニール容器", "ビン", "プラ容器(ヤクルト型)", "プラ容器(ヨーグルト型)", "ペットボトル", "ペットボトル(ミニ)", "ボトル缶", "ポーション", "箱(飲料)", "紙パック(Pキャン)", "紙パック(キャップ付き)", "紙パック(ゲーブルトップ)", "紙パック(ブリックパック)", "紙パック(円柱型)", "缶"] # 获取字典格式的报告 report = classification_report(y_true, y_pred, target_names=class_names, output_dict=True) # 自定义列宽:类名列自动适配最长名称+4字符间距,其余列固定宽度,可按需调整数值 name_col_width = max([len(name) for name in class_names]) + 4 col_config = [ ("", name_col_width), ("precision", 12), ("recall", 10), ("f1-score", 10), ("support", 10) ] # 拼接输出内容 output_lines = [] # 拼接表头 output_lines.append("".join([header.ljust(width) for header, width in col_config])) output_lines.append("") # 拼接每个类别的指标行 for cls_name in class_names: metrics = report[cls_name] row = [ cls_name.ljust(col_config[0][1]), f"{metrics['precision']:.2f}".ljust(col_config[1][1]), f"{metrics['recall']:.2f}".ljust(col_config[2][1]), f"{metrics['f1-score']:.2f}".ljust(col_config[3][1]), f"{int(metrics['support'])}".ljust(col_config[4][1]) ] output_lines.append("".join(row)) output_lines.append("") # 拼接汇总行 total_support = int(report["macro avg"]["support"]) output_lines.append( "accuracy".ljust(col_config[0][1]) + "".ljust(col_config[1][1] + col_config[2][1]) + f"{report['accuracy']:.2f}".ljust(col_config[3][1]) + f"{total_support}".ljust(col_config[4][1]) ) for avg_type in ["macro avg", "weighted avg"]: metrics = report[avg_type] row = [ avg_type.ljust(col_config[0][1]), f"{metrics['precision']:.2f}".ljust(col_config[1][1]), f"{metrics['recall']:.2f}".ljust(col_config[2][1]), f"{metrics['f1-score']:.2f}".ljust(col_config[3][1]), f"{int(metrics['support'])}".ljust(col_config[4][1]) ] output_lines.append("".join(row)) # 生成最终格式化报告,可直接打印或保存为txt final_report = "\n".join(output_lines) print(final_report) with open("aligned_classification_report.txt", "w", encoding="utf-8") as f: f.write(final_report)
只需要调整col_config里的宽度数值,就能自由控制各列之间的间距,适配任意长度的类名。
方案2:补丁修改sklearn内部宽度计算逻辑
如果不想重写全量格式化逻辑,可以临时修改sklearn生成报告时的宽度计算规则,手动给类名列增加预留宽度:
from sklearn.metrics import classification_report import sklearn.metrics._classification as clf_module # 保存原方法 original_check = clf_module._check_targets # 定义补丁方法,增加类名预留宽度 def patched_check(y_true, y_pred, *, target_names=None): y_type, y_true, y_pred = original_check(y_true, y_pred) if target_names is not None: # 额外增加6字符的类名列宽度,按需调整数值即可 clf_module._width = max([len(name) for name in target_names]) + 6 return y_type, y_true, y_pred # 替换原方法 clf_module._check_targets = patched_check # 后续正常调用classification_report即可,生成的字符串会自动加宽类名列 report_str = classification_report(y_true, y_pred, target_names=class_names) with open("report.txt", "w", encoding="utf-8") as f: f.write(report_str)
注意:该方法依赖sklearn内部私有函数,不同版本的sklearn内部函数名可能存在差异,兼容性不如方案1。
方案3:显示层优化
如果不需要修改代码,保存txt文件后,使用VS Code、Notepad++等编辑器打开,将显示字体设置为等宽字体(如更纱黑体、MS Gothic),可以缓解全角字符导致的错位问题,但该方法没有从根源调整列间距,适配性有限。
内容的提问来源于stack exchange,提问作者sksoumik
相关产品推荐
相关产品推荐

