You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Orange3脚本调用RandomForestLearner无法复现GUI结果求助

Fixing "ValueError: Input contains NaN..." in Orange3 Random Forest Script

Hey there! I’ve run into this exact issue before—Orange’s GUI does a lot of behind-the-scenes data cleaning that doesn’t happen automatically when you use the scripting API. Let’s break down why this is happening and how to fix it.

Why the GUI Works but Your Script Doesn’t

Orange’s GUI workflows automatically handle common data issues like missing values (NaN) when you load data through its components. But when you use Orange.data.Table directly in a script, it doesn’t apply these preprocessing steps by default. The Random Forest learner (and most Orange learners) will throw that NaN error if the input data has missing values or invalid entries.

Step 1: Verify Missing Values

First, confirm that your datasets actually have missing values. Add these lines to your script to check:

import pandas as pd

# Convert Orange Table to pandas DataFrame for easier inspection
train_df = pd.DataFrame(train.X, columns=train.domain.attributes)
test_df = pd.DataFrame(test.X, columns=test.domain.attributes)

print(f"Total missing values in train set: {train_df.isnull().sum().sum()}")
print(f"Total missing values in test set: {test_df.isnull().sum().sum()}")

Step 2: Impute Missing Values (Match GUI Behavior)

Orange has a built-in Imputer preprocessor that mimics the GUI’s default behavior—filling numerical columns with the mean, categorical columns with the mode. Here’s how to integrate it into your script:

import os
import Orange
from Orange.preprocess import Imputer

cwd = os.getcwd() + '\\'
train = Orange.data.Table(cwd + 'train_ex.csv')
test = Orange.data.Table(cwd + 'test_ex.csv')

# Clean data with imputation
imputer = Imputer()
train_clean = imputer(train)
test_clean = imputer(test)

# Train and evaluate the model
learner = Orange.classification.RandomForestLearner()
result = Orange.evaluation.testing.TestOnTestData(train_clean, test_clean, [learner])

# Calculate and print metrics
MSE = Orange.evaluation.MSE(result)
RMSE = Orange.evaluation.RMSE(result)
MAE = Orange.evaluation.MAE(result)
R2 = Orange.evaluation.R2(result)

print(f"MSE: {MSE[0]}, RMSE: {RMSE[0]}, MAE: {MAE[0]}, R2: {R2[0]}")

Bonus: Check for Infinite Values

If you still get the error after imputing, check for infinite values in your data:

print(f"Total infinite values in train set: {train_df.isin([float('inf'), -float('inf')]).sum().sum()}")
print(f"Total infinite values in test set: {test_df.isin([float('inf'), -float('inf')]).sum().sum()}")

If you find any, replace them with column means (then convert back to an Orange Table):

train_df.replace([float('inf'), -float('inf')], train_df.mean(), inplace=True)
train_clean = Orange.data.Table.from_frame(train_df, domain=train.domain)

test_df.replace([float('inf'), -float('inf')], test_df.mean(), inplace=True)
test_clean = Orange.data.Table.from_frame(test_df, domain=test.domain)

This should replicate the same behavior you see in the Orange GUI and get your random forest model running without errors.

内容的提问来源于stack exchange,提问作者Josh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:09:13