传感器值校正回归算法选型及2011年1月实际值预测方法咨询
Hey, let's work through how to predict those actual values for January 2011. Here's a step-by-step approach tailored to your data and questions:
Your data has three key traits: it's multi-dimensional (different application + sensor combinations), has missing monthly entries, and includes both categorical (application/sensor) and numerical (measured, time) features. Tree-based models like CatBoost or XGBoost are the best fit here—they handle categorical data natively, capture non-linear relationships between sensor bias and other features, and work well with incomplete time-series groups.
1. First: Clean Your Data
Your raw data uses "." as a missing value marker, and the time column isn't in a machine-readable format. Let's fix that first:
library(tidyverse) library(lubridate) library(catboost) # Clean training data train_clean <- train_data %>% # Replace "." with proper NA values and convert numeric columns mutate( across(c(measured, actual), ~ifelse(. == ".", NA_real_, as.numeric(.))), # Convert time to a date object (day set to 01 since it's monthly data) time = dmy(paste0("01.", time)), # Extract year/month as separate features (easier for models to use than raw dates) year = year(time), month = month(time) ) %>% # Remove rows with missing time values (the ones marked ".") filter(!is.na(time)) # Clean test data to match training data structure test_clean <- test_data %>% mutate( measured = as.numeric(measured), time = dmy("01.01.2011"), year = year(time), month = month(time) )
2. Model Selection: Why Tree Models Beat ARIMA/Neural Networks
Let's break down why your initial ideas work (or don't):
- ARIMA: Not ideal here. ARIMA is designed for single, continuous time series. Your data has dozens of disjoint time series (one per
application+sensor), many with missing months—you'd have to train a separate ARIMA model for each group, which is inefficient and prone to poor performance with small data. - Neural Networks: 10 years of monthly data gives ~120 points total, but most
application+sensorgroups have far fewer. Neural networks need large datasets to generalize well, so you'd likely end up overfitting to noise. - CatBoost/XGBoost: Perfect fit. They:
- Handle categorical features (
application,sensor) without manual one-hot encoding (CatBoost does this automatically) - Capture non-linear patterns in sensor bias (like how
factorysensors might have consistent bias vsresearchsensors) - Work with incomplete time series groups
- Handle categorical features (
You tried using diff = measured - actual as a feature—this is a great idea! The bias between measured and actual values is likely tied to the sensor type, application, and time, so including diff as a feature (or even predicting diff first, then calculating actual = measured - predicted_diff) will help the model learn those patterns.
3. Step-by-Step Implementation with CatBoost
CatBoost is easier to use for categorical data, so we'll use it here.
Add the diff Feature
First, calculate the bias column for training data:
train_clean <- train_clean %>% mutate(diff = measured - actual)
Prepare Data for CatBoost
CatBoost uses "pools" to handle data, and we need to flag categorical features:
# Define features: exclude raw time and target (actual) feature_cols <- c("application", "sensor", "measured", "year", "month", "diff") target_col <- "actual" # Create training pool train_pool <- catboost.load_pool( data = train_clean %>% select(all_of(feature_cols)), label = train_clean[[target_col]], cat_features = c("application", "sensor") # Tell CatBoost these are categorical ) # Create test pool (note: test data doesn't have diff, so we set it to NA; CatBoost handles this) test_pool <- catboost.load_pool( data = test_clean %>% mutate(diff = NA) %>% select(all_of(feature_cols)), cat_features = c("application", "sensor") )
Train the Model with Cross-Validation
We'll use cross-validation to avoid overfitting and find the optimal number of iterations:
# Set model parameters params <- list( iterations = 1000, learning_rate = 0.05, depth = 6, loss_function = "RMSE", eval_metric = "RMSE", early_stopping_rounds = 50, # Stop if validation RMSE doesn't improve for 50 rounds verbose = 100 # Print progress every 100 iterations ) # Run cross-validation to find best iteration count cv_result <- catboost.cv(train_pool, params, fold_count = 5) best_iter <- which.min(cv_result$test.RMSE.mean) # Train final model with optimal iterations model <- catboost.train(train_pool, params = modifyList(params, list(iterations = best_iter)))
Predict on Test Data
Finally, generate predictions for January 2011:
# Generate predictions test_clean$predicted_actual <- catboost.predict(model, test_pool) # View results test_clean %>% select(application, sensor, measured, predicted_actual)
4. Key Tips for Improvement
- Time-Based Validation: Instead of random cross-validation, use time sliding windows (e.g., train on 2000-2008, validate on 2009; train on 2000-2009, validate on 2010) to mimic real-world forecasting.
- Group-Specific Features: Add features like the average
difffor eachapplication+sensorgroup, or rolling averages ofdiffover past months (handle missing values withna.rm = TRUE). - Alternative Target: Instead of predicting
actualdirectly, predictdifffirst, then computeactual = measured - predicted_diff. This can sometimes lead to better performance since the model focuses on learning the bias pattern.
内容的提问来源于stack exchange,提问作者Abhishek

