You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

构建random forest regression时出现object not found错误求助

解决Shiny应用中“object 'offense_category' not found”错误及模型预测问题

问题根源

  1. 测试数据格式完全错误:你在server端构建测试数据时,错误地将特征列转成了Name-Value的长格式,加上冗余的写/读csv操作,导致最终的test数据框根本没有模型所需的offense_category、year、neighborhood列名,模型预测时自然找不到对应特征。
  2. 全局数据加载缺失:UI中直接引用df$offense_category,如果df未在全局环境提前加载,会导致UI初始化时找不到该对象。
  3. 模型类型混淆:你提到用随机森林回归,但预测时调用type="prob"——该参数仅适用于分类模型,回归模型不支持此参数。预测逮捕罪名属于分类任务,需确保模型为分类类型。

修复步骤及代码示例

1. 全局数据与模型加载

创建global.R文件(或在app开头执行),提前加载数据集和训练模型到全局环境:

# global.R
library(randomForest)
library(shiny)
library(shinythemes)

# 加载你的犯罪数据集(替换为实际路径)
df <- read.csv("your_crime_dataset.csv")
# 划分训练集(示例,需根据实际需求调整)
trainingset <- df

# 将目标变量转为因子(分类模型要求)
trainingset$charge <- as.factor(trainingset$charge)
# 构建随机森林分类模型
rf <- randomForest(charge ~ offense_category+year+neighborhood, 
                   data=trainingset, ntree=500, importance=TRUE)

2. 修复UI代码

确保下拉选项与训练集类别一致,避免依赖未加载的全局变量:

# ui.R
ui <- fluidPage(theme = shinytheme("united"),
                
                headerPanel("Prediction"),
                sidebarPanel(
                  h2("Filter"),
                  selectizeInput("offense_category", label="Offense Category", 
                                 choices=c("", sort(unique(trainingset$offense_category))),
                                 multiple=FALSE),
                  radioButtons(
                    "year",
                    "Select year",
                    choices = c("2020", "2021","2022","2023"),
                    selected = "2020"
                  ),
                  selectizeInput("neighborhood", label="Neighborhood", 
                                 choices=c("", sort(unique(trainingset$neighborhood))),
                                 multiple=FALSE),
                  actionButton("predictbutton", "Predict", class = "btn btn-primary")
                ),
                mainPanel(h2("Prediction"), 
                          tableOutput(outputId = 'chargeprediction'))
)

3. 重构Server端预测逻辑

去掉冗余的文件操作,直接构建符合模型要求的测试数据,并处理类型匹配:

# server.R
server <- function(input, output){
  datainput <- reactive({
    # 验证用户输入非空
    validate(
      need(input$offense_category != "", "请选择犯罪类别"),
      need(input$neighborhood != "", "请选择社区")
    )
    
    # 构建模型所需格式的测试数据,确保因子水平与训练集一致
    test <- data.frame(
      offense_category = factor(input$offense_category, levels = levels(trainingset$offense_category)),
      year = factor(input$year, levels = levels(trainingset$year)),
      neighborhood = factor(input$neighborhood, levels = levels(trainingset$neighborhood)),
      stringsAsFactors = FALSE
    )
    
    # 生成预测结果
    pred_class <- predict(rf, test)
    pred_prob <- predict(rf, test, type="prob")
    
    # 整理输出表格
    data.frame(
      Prediction = pred_class,
      round(pred_prob, 1)
    )
  })
  
  output$chargeprediction <- renderTable({
    if (input$predictbutton > 0){
      isolate(datainput())
    }
  })
}

额外注意事项

  • 确保训练集与测试集的特征类型完全匹配(如因子水平、数值/因子类型),否则模型会出现类型不兼容错误。
  • 如果year在训练集中是数值类型,需将测试集的input$year转为数值(as.numeric(input$year)),而非因子。

内容的提问来源于stack exchange,提问作者Chivosner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 06:43:18