You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Weka(Java)中让多层感知器预测同时输出字符串类别与概率

问题描述

使用Multilayer Perceptron(MLP)做预测时,已生成待预测的测试数据,通过以下代码遍历记录并追加预测结果:

for (int i1 = 0; i1 < datapredict1.numInstances(); i1++) {      
            double clsLabel1 = mlp.classifyInstance(datapredict1.instance(i1));
            datapredict1.instance(i1).setClassValue(clsLabel1); 
            String s = datapredict1.instance(i1) + "," + clsLabel1;
            writer11.write(s.toString());
            writer11.newLine();
            System.out.println(datapredict1.instance(i1) + "," + clsLabel1);
        }

当前输出的类别为数值形式,且重复输出了索引值:

0.178571,0.2,0.181818,0.333333,0,09:15,0.849899,0.8498991728827364
0.414835,0,0.454545,0.666667,0,16:15,0.850662,0.85066198399766

期望输出为带引号的字符串类别值,同时保留预测概率:

0.178571,0.2,0.181818,0.333333,0,09:15,"Value2",0.8498991728827364
0.414835,0,0.454545,0.666667,0,16:15,"Value4",0.85066198399766

解决方案

要实现需求,需要从Weka的类别属性中获取索引对应的字符串值,并正确提取预测概率,修改后的代码如下:

// 提前获取数据集的类别属性,避免循环内重复调用
Attribute classAttr = datapredict1.classAttribute();

for (int i1 = 0; i1 < datapredict1.numInstances(); i1++) {      
    Instance inst = datapredict1.instance(i1);
    // 获取预测的类别索引
    double clsIndex = mlp.classifyInstance(inst);
    // 将索引转换为对应的字符串类别值
    String clsStr = classAttr.value((int) clsIndex);
    // 获取该类别对应的预测概率
    double[] probDist = mlp.distributionForInstance(inst);
    double clsProb = probDist[(int) clsIndex];
    
    // 手动构建输出行:拼接所有非类别属性,再追加带引号的类别和概率
    StringBuilder sb = new StringBuilder();
    for (int j = 0; j < inst.numAttributes(); j++) {
        if (j == inst.classIndex()) {
            continue; // 跳过原数据集中的类别列
        }
        if (sb.length() > 0) {
            sb.append(",");
        }
        sb.append(inst.value(j));
    }
    // 追加带引号的类别字符串和概率
    sb.append(",\"").append(clsStr).append("\",").append(clsProb);
    
    String outputLine = sb.toString();
    writer11.write(outputLine);
    writer11.newLine();
    System.out.println(outputLine);
}

关键说明

  1. 类别索引转字符串:通过classAttribute().value((int) clsIndex)将MLP返回的类别索引(数值)转换为数据集定义的字符串类别名。
  2. 获取预测概率:使用distributionForInstance()方法获取所有类别的概率分布,再取出对应预测类别的概率,替代原代码中重复输出的索引值。
  3. 手动构建输出行:跳过原数据集中的类别列,避免输出数值形式的类别,替换为带引号的字符串类别值,保证输出格式符合预期。

内容的提问来源于stack exchange,提问作者Micmac

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 09:36:10