You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言XGBoost时间序列预测中mutate()报错及代码修复方案

Got it, let's work through fixing those errors and getting your supplier-level XGBoost time series predictions sorted. I'll walk through the exact issues you're hitting (date parsing, missing objects, missing columns) and provide a fully reproducible, corrected workflow.

1. First: Fix Data & Date Parsing Errors

The mutate() date error almost always happens because your date column isn't a proper Date type. Let's start by creating a reproducible mydat dataset (matching your use case) and ensuring dates are formatted correctly.

# Generate a reproducible mydat dataset (mimicking your structure)
set.seed(123)
suppliers <- c("Supplier_A", "Supplier_B", "Supplier_C")
dates <- seq(as.Date("2023-01-01"), as.Date("2023-12-31"), by = "day")

mydat <- expand.grid(supplier = suppliers, date = dates) %>%
  mutate(
    # Simulate base prices with weekday trends and slow growth
    base_price = rnorm(nrow(.), mean = 50, sd = 5) + 
      ifelse(lubridate::wday(date) %in% c(6,7), 3, 0) +
      as.numeric(date - min(date))*0.01,
    # Convert weekday to numeric (XGBoost needs numerical features)
    weekday = lubridate::wday(date, label = FALSE)
  )

If your original mydat has dates stored as strings, fix them with:

# Fix string-to-date conversion (adjust format to match your data)
mydat <- mydat %>% mutate(date = as.Date(date, format = "%Y-%m-%d"))
2. Fix "Object Not Found" & "Column Doesn't Exist" Errors

These usually come from:

  • Typos in column names (e.g., baseprice instead of base_price)
  • Not properly referencing grouped data
  • Missing time-series features (XGBoost isn't a native time-series model—you need to add lag/rolling features manually)

Here's the corrected grouped modeling workflow:

library(tidyverse)
library(lubridate)
library(xgboost)
library(zoo)

# Step 1: Preprocess data + create time-series features
mydat_processed <- mydat %>%
  mutate(
    # Ensure date is strictly Date type
    date = as.Date(date),
    # Add lag features (past 1 and 7 days' prices)
    lag1 = lag(base_price, 1),
    lag7 = lag(base_price, 7),
    # Add 7-day rolling average price
    roll_mean7 = zoo::rollmean(base_price, k=7, fill=NA, align="right")
  ) %>%
  drop_na()  # Remove rows with missing features (avoids modeling errors)

# Step 2: Group by supplier, train XGBoost, and predict next 7 days
supplier_forecasts <- mydat_processed %>%
  group_by(supplier) %>%
  group_modify(function(group_data, supplier_key) {
    # Define features (X) and target (y) for the group
    X_features <- group_data %>% select(weekday, lag1, lag7, roll_mean7) %>% as.matrix()
    y_target <- group_data$base_price
    
    # Train XGBoost model
    xgb_model <- xgboost(
      data = X_features,
      label = y_target,
      nrounds = 100,
      objective = "reg:squarederror",
      verbose = 0  # Mute training logs
    )
    
    # Create future 7-day data for prediction
    last_date <- max(group_data$date)
    future_dates <- seq(last_date + 1, last_date + 7, by = "day")
    
    future_features <- tibble(
      date = future_dates,
      weekday = wday(future_dates, label = FALSE),
      # Use latest known values for lag features
      lag1 = last(group_data$base_price),
      lag7 = group_data$base_price[nrow(group_data)-6],  # Price from same weekday last week
      roll_mean7 = last(group_data$roll_mean7)
    )
    
    # Generate predictions
    future_X <- future_features %>% select(weekday, lag1, lag7, roll_mean7) %>% as.matrix()
    future_features$predicted_base_price <- predict(xgb_model, future_X)
    
    return(future_features)
  }) %>%
  ungroup() %>%
  # Reorder columns to match your target format
  select(supplier, date, weekday, predicted_base_price)

# Check the final forecast output
head(supplier_forecasts)
3. Key Fixes Explained
  • Date parsing error: We explicitly convert date to Date type with as.Date(), and specify the format if your original dates are strings. This eliminates mutate() errors related to date handling.
  • Object/column not found:
    • We use group_data inside group_modify() to reference the current supplier's subset of data (instead of the global mydat).
    • All column names are consistent (e.g., base_price instead of typos), and we only select columns that exist in the processed data.
  • XGBoost for time series: We added critical time-series features (lag1, lag7, roll_mean7) because XGBoost can't inherently understand time order. Combined with weekday, these features let the model learn weekly patterns and trends.

内容的提问来源于stack exchange,提问作者psysky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 20:17:29