关于R包Biocomb中FCFS的attrs.nominal参数及报错的技术咨询
Hi there! Let's break down your two questions step by step, since you're working through feature selection with Biocomb's FCFS and hitting a couple of roadblocks.
Understanding the
attrs.nominal Parameter in Biocomb's FCFS First, let's clarify what this parameter does and how to use it correctly:
- Core purpose: FCFS (Fast Correlation Filter Selection) uses different statistical metrics to calculate feature relevance and redundancy depending on whether variables are continuous or nominal (categorical). The
attrs.nominalparameter explicitly tells Biocomb which columns in your dataset are nominal, so it can apply the appropriate correlation measures (like mutual information for categorical pairs, instead of Pearson correlation for continuous ones). - Do you need to include all nominal variables?: Yes, you absolutely should. If you omit a nominal variable from this parameter, Biocomb will default to treating it as continuous, which will skew your correlation calculations and lead to inaccurate feature selection results.
- How to specify it: You can pass either column names (as a character vector) or column indices (as an integer vector). For example:
Double-check that the columns you list are truly categorical (non-continuous) in your dataset.# Using column names for categorical variables fcfs_result <- FCFS(data = train_data, class = "target_label", attrs.nominal = c("gender", "disease_stage")) # Using column indices for categorical variables fcfs_result <- FCFS(data = train_data, class = "target_label", attrs.nominal = c(2, 5))
Fixing the "Undefined Columns Selected" Error
That error (Error in [.data.frame (data.validation, , 72) : undefined columns selected) is a common one—it simply means your code is trying to access column 72 in data.validation, but this dataset doesn't have 72 columns. Here's how to troubleshoot and fix it:
- Verify your validation dataset's column count: Run
ncol(data.validation)to see how many columns are actually present. If the number is less than 72, you've found the root cause immediately. - Trace where column 72 is referenced: Look through your code for lines like
data.validation[, 72]or any subsetting that uses index 72. Common issues include:- Accidentally hardcoding a column index that only exists in your training set (not the validation set)
- A mismatch between the feature indices returned by FCFS and the validation set's column structure
- Ensure consistent column structure between train and validation sets: The feature indices FCFS returns correspond to columns in your training data. Make sure your validation set has the exact same feature columns (in the same order) as the training set. If you modified either dataset (e.g., dropped columns, reordered them), this will cause mismatches.
- Safe subsetting with selected features: After running FCFS, use the returned selected feature indices/names to subset both datasets properly. This avoids hardcoding indices entirely. Example:
# Extract selected features from FCFS result selected_attrs <- fcfs_result$selected_attrs # Subset training and validation data to only include these features train_selected <- train_data[, selected_attrs] validation_selected <- data.validation[, selected_attrs] # Proceed with SVM training and testing svm_model <- svm(target_label ~ ., data = train_selected) svm_pred <- predict(svm_model, newdata = validation_selected)
内容的提问来源于stack exchange,提问作者Aymen Trabelsi
相关产品推荐
相关产品推荐

