You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用partykit包varimp计算条件重要性时报错求助

Error: NAs are not allowed in subscripted assignments when using varimp(cforest, conditional = TRUE) with stratified sampling

Problem Description

I built a cforest model to handle a highly imbalanced dataset (only 7% of samples have y=1), using the strata parameter to ensure minority class samples aren't ignored during bagging:

cf1 <- cforest(y~., data = DATA, strata = DATA$y, ntree = 200L, mtry = 10)

The model's confusion matrix showed it was working correctly, but when I tried to calculate conditional variable importance:

cf1.imp_cond <- varimp(cf1, conditional = TRUE)

I got this error:

Error in x[strata == s] <- .resample(x[strata == s]) : NAs are not allowed in subscripted assignments

I traced the error to the kidids_node function call, and was able to reproduce the issue with test data using this code:

cf2 <- cforest(X5_years_survival~., data = test, strata = X5_years_survival, ntree = 200L, mtry = 6)
cf2.imp_cond <- varimp(cf2, conditional = TRUE)

Same error message every time. Has anyone run into this issue before, and how can I fix it?


Solutions & Troubleshooting Steps

This error typically happens when stratified sampling leads to missing strata in some tree nodes, or there are NA values in your stratification variable—a common pain point with imbalanced data and conditional variable importance calculations. Here are actionable fixes:

  1. Check for NA values in your stratification variable
    First, verify if your target/strata column has any missing values. Even though cforest handles missing features, NA labels in the stratification variable break the resampling step for conditional importance.

    # Check for NAs in the strata column
    sum(is.na(DATA$y))
    

    If there are NAs, clean your data first:

    DATA_clean <- na.omit(DATA)
    cf1 <- cforest(y~., data = DATA_clean, strata = DATA_clean$y, ntree = 200L, mtry = 10)
    
  2. Disable within-strata permutation in varimp
    By default, varimp uses permute.within.strata = TRUE when the model was built with strata. But if some nodes have no samples from a particular stratum (super common with imbalanced data), this triggers the NA subscript error. Turn off this parameter:

    cf1.imp_cond <- varimp(cf1, conditional = TRUE, permute.within.strata = FALSE)
    
  3. Stabilize your forest structure
    A small number of trees (ntree=200) can lead to unstable nodes with tiny sample sizes, where entire strata might be missing. Try increasing the number of trees and adjusting mtry to improve node stability:

    # Use a larger ntree and default mtry (sqrt of feature count)
    cf1 <- cforest(y~., data = DATA, strata = DATA$y, ntree = 500L, mtry = floor(sqrt(ncol(DATA)-1)))
    
  4. Fallback to unconditional variable importance (temp workaround)
    If the above steps don't work immediately, you can use unconditional importance for initial analysis while troubleshooting:

    cf1.imp_uncond <- varimp(cf1, conditional = FALSE)
    

Why This Happens

Conditional variable importance works by permuting each feature within individual tree nodes. When you use stratified sampling, the function tries to permute only within each stratum to preserve the class distribution. But if a node has no samples from a stratum (e.g., no minority class samples), the check strata == s returns NA for those positions, causing the assignment error you see.

内容的提问来源于stack exchange,提问作者Bs He

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:57:35