如何获取xgb.DMatrix索引或添加信息以用于feval函数调参?
Great question—both approaches you’ve outlined are valid and effective ways to associate additional sample information with your custom feval function in XGBoost for R. Let’s break down how each method works and provide complete implementations:
1. Attaching Custom Metadata Directly to xgb.DMatrix
This is the most straightforward approach, exactly as you’ve started. You can store any custom data (like your ID field) directly in the xgb.DMatrix object during creation, then retrieve it within your feval function using getinfo().
Full Implementation Example
library(xgboost) library(data.table) # Sample input data OBSERVATIONS <- data.table( actual = rnorm(100), # Target values ID = paste0("ID_", 1:100) # Unique sample IDs ) scalingFact <- data.table( ID = paste0("ID_", 1:100), # Matching IDs scale = runif(100, 0.8, 1.2) # Extra metadata to use in evaluation ) feature_matrix <- matrix(rnorm(100 * 10), nrow = 100) # Model features # Create DMatrix with custom ID metadata dtrain <- xgb.DMatrix( data = feature_matrix, label = OBSERVATIONS$actual, ID = OBSERVATIONS$ID # Attach custom field ) # Complete custom evaluation function custom_feval <- function(preds, dtrain) { # Extract label and custom metadata from DMatrix actual_values <- getinfo(dtrain, "label") sample_ids <- getinfo(dtrain, "ID") # Combine predictions, actuals, and IDs into a data table eval_data <- data.table( ID = sample_ids, actual = actual_values, preds = preds ) # Join with your additional metadata eval_data <- scalingFact[eval_data, on = "ID"] # Calculate your custom metric (example: scaled MAE) scaled_mean_abs_error <- mean(abs(eval_data$actual - eval_data$preds) * eval_data$scale) # Return metric in the format XGBoost expects return(list(metric = "scaled_mae", value = scaled_mean_abs_error)) } # Test with cross-validation xgb.cv( data = dtrain, nrounds = 10, nfold = 5, feval = custom_feval, verbose = 1 )
2. Accessing Sample Indices from xgb.DMatrix
You can also retrieve the indices of samples in the xgb.DMatrix to link back to your original dataset. By default, XGBoost assigns indices starting at 0, but you can explicitly set custom indices (like your original data’s row numbers) when creating the DMatrix for clearer alignment.
Full Implementation Example
# Create DMatrix with explicit sample indices (using original row numbers) dtrain <- xgb.DMatrix( data = feature_matrix, label = OBSERVATIONS$actual, index = 1:nrow(OBSERVATIONS) # Use original data's row indices ) # Custom evaluation function using indices custom_feval_with_indices <- function(preds, dtrain) { # Extract indices and target values from DMatrix sample_indices <- getinfo(dtrain, "index") actual_values <- getinfo(dtrain, "label") # Fetch sample IDs and other metadata using indices from original data eval_data <- OBSERVATIONS[sample_indices, .(ID, actual = actual_values)] eval_data$preds <- preds # Join with additional metadata eval_data <- scalingFact[eval_data, on = "ID"] # Calculate custom metric scaled_mean_abs_error <- mean(abs(eval_data$actual - eval_data$preds) * eval_data$scale) return(list(metric = "scaled_mae", value = scaled_mean_abs_error)) } # Test the function xgb.cv( data = dtrain, nrounds = 10, nfold = 5, feval = custom_feval_with_indices, verbose = 1 )
Which Approach Should You Use?
- Use the custom metadata method if you want to store the exact information you need directly in the DMatrix—this avoids relying on external data row order and makes your
fevalfunction self-contained. - Use the index method if you prefer to keep metadata in a separate dataset and just need a way to map predictions back to their original samples.
Both methods fully meet your requirement of linking extra observation information to your custom evaluation logic.
内容的提问来源于stack exchange,提问作者P_Sta

