You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

传感器值校正回归算法选型及2011年1月实际值预测方法咨询

Hey, let's work through how to predict those actual values for January 2011. Here's a step-by-step approach tailored to your data and questions:

Approach Overview

Your data has three key traits: it's multi-dimensional (different application + sensor combinations), has missing monthly entries, and includes both categorical (application/sensor) and numerical (measured, time) features. Tree-based models like CatBoost or XGBoost are the best fit here—they handle categorical data natively, capture non-linear relationships between sensor bias and other features, and work well with incomplete time-series groups.

1. First: Clean Your Data

Your raw data uses "." as a missing value marker, and the time column isn't in a machine-readable format. Let's fix that first:

library(tidyverse)
library(lubridate)
library(catboost)

# Clean training data
train_clean <- train_data %>%
  # Replace "." with proper NA values and convert numeric columns
  mutate(
    across(c(measured, actual), ~ifelse(. == ".", NA_real_, as.numeric(.))),
    # Convert time to a date object (day set to 01 since it's monthly data)
    time = dmy(paste0("01.", time)),
    # Extract year/month as separate features (easier for models to use than raw dates)
    year = year(time),
    month = month(time)
  ) %>%
  # Remove rows with missing time values (the ones marked ".")
  filter(!is.na(time))

# Clean test data to match training data structure
test_clean <- test_data %>%
  mutate(
    measured = as.numeric(measured),
    time = dmy("01.01.2011"),
    year = year(time),
    month = month(time)
  )

2. Model Selection: Why Tree Models Beat ARIMA/Neural Networks

Let's break down why your initial ideas work (or don't):

  • ARIMA: Not ideal here. ARIMA is designed for single, continuous time series. Your data has dozens of disjoint time series (one per application+sensor), many with missing months—you'd have to train a separate ARIMA model for each group, which is inefficient and prone to poor performance with small data.
  • Neural Networks: 10 years of monthly data gives ~120 points total, but most application+sensor groups have far fewer. Neural networks need large datasets to generalize well, so you'd likely end up overfitting to noise.
  • CatBoost/XGBoost: Perfect fit. They:
    • Handle categorical features (application, sensor) without manual one-hot encoding (CatBoost does this automatically)
    • Capture non-linear patterns in sensor bias (like how factory sensors might have consistent bias vs research sensors)
    • Work with incomplete time series groups

You tried using diff = measured - actual as a feature—this is a great idea! The bias between measured and actual values is likely tied to the sensor type, application, and time, so including diff as a feature (or even predicting diff first, then calculating actual = measured - predicted_diff) will help the model learn those patterns.

3. Step-by-Step Implementation with CatBoost

CatBoost is easier to use for categorical data, so we'll use it here.

Add the diff Feature

First, calculate the bias column for training data:

train_clean <- train_clean %>%
  mutate(diff = measured - actual)

Prepare Data for CatBoost

CatBoost uses "pools" to handle data, and we need to flag categorical features:

# Define features: exclude raw time and target (actual)
feature_cols <- c("application", "sensor", "measured", "year", "month", "diff")
target_col <- "actual"

# Create training pool
train_pool <- catboost.load_pool(
  data = train_clean %>% select(all_of(feature_cols)),
  label = train_clean[[target_col]],
  cat_features = c("application", "sensor") # Tell CatBoost these are categorical
)

# Create test pool (note: test data doesn't have diff, so we set it to NA; CatBoost handles this)
test_pool <- catboost.load_pool(
  data = test_clean %>% mutate(diff = NA) %>% select(all_of(feature_cols)),
  cat_features = c("application", "sensor")
)

Train the Model with Cross-Validation

We'll use cross-validation to avoid overfitting and find the optimal number of iterations:

# Set model parameters
params <- list(
  iterations = 1000,
  learning_rate = 0.05,
  depth = 6,
  loss_function = "RMSE",
  eval_metric = "RMSE",
  early_stopping_rounds = 50, # Stop if validation RMSE doesn't improve for 50 rounds
  verbose = 100 # Print progress every 100 iterations
)

# Run cross-validation to find best iteration count
cv_result <- catboost.cv(train_pool, params, fold_count = 5)
best_iter <- which.min(cv_result$test.RMSE.mean)

# Train final model with optimal iterations
model <- catboost.train(train_pool, params = modifyList(params, list(iterations = best_iter)))

Predict on Test Data

Finally, generate predictions for January 2011:

# Generate predictions
test_clean$predicted_actual <- catboost.predict(model, test_pool)

# View results
test_clean %>% select(application, sensor, measured, predicted_actual)

4. Key Tips for Improvement

  • Time-Based Validation: Instead of random cross-validation, use time sliding windows (e.g., train on 2000-2008, validate on 2009; train on 2000-2009, validate on 2010) to mimic real-world forecasting.
  • Group-Specific Features: Add features like the average diff for each application+sensor group, or rolling averages of diff over past months (handle missing values with na.rm = TRUE).
  • Alternative Target: Instead of predicting actual directly, predict diff first, then compute actual = measured - predicted_diff. This can sometimes lead to better performance since the model focuses on learning the bias pattern.

内容的提问来源于stack exchange,提问作者Abhishek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:21:15