R语言训练kNN模型触发na.fail.default错误,求解决方法
Hey there! Let's break down why you're hitting that na.fail.default error and how to fix it. The core issue here is how you're converting your term-document matrix (TDM) into a data frame—using inspect() is causing all those NA values. Let's walk through the solution step by step.
Why You're Seeing This Error
When you run:
tdm <- as.data.frame(inspect(tdm)) weightedtdm <- as.data.frame(inspect(weightedtdm))
You're making a common beginner mistake. The inspect() function is meant only for viewing the content of your TDM—it outputs a human-readable formatted text, not a structured matrix. Converting that output directly to a data frame confuses R, leading to a data frame filled with NA values. When you split this into tdmTrain, you end up with a dataset full of missing values, which the train() function can't handle for kNN.
Step-by-Step Fixes
1. Correctly Convert TDM to Data Frame
Replace the lines where you use inspect() with code that first converts the TDM to a matrix, then to a data frame. This preserves the structured numerical data:
# Convert tdm's into data frames correctly tdm <- as.data.frame(as.matrix(tdm)) weightedtdm <- as.data.frame(as.matrix(weightedtdm))
This ensures your TDM data is properly formatted as numerical values instead of messy NA-filled text.
2. (Optional) Filter Empty Documents
Another potential issue could be empty documents after preprocessing (e.g., a document that had only stopwords, numbers, or punctuation). Empty documents can cause unexpected behavior, so add this code after preprocessing your corpus to filter them out:
# Remove empty documents from corpus and merged data frame empty_docs <- sapply(corpus, function(x) content(x) == "") corpus <- corpus[!empty_docs] merged <- merged[!empty_docs, ]
This keeps your dataset consistent and avoids any edge cases with empty text.
3. Verify No More NA Values
After making these changes, check if your training data has any NA values before fitting the model:
# Check for NA values in weightedTDMtrain any(is.na(weightedTDMtrain))
If this returns FALSE, you're good to go!
4. Re-Run Your kNN Model
Now your weightedTDMtrain should be a clean, NA-free dataset. Re-run your train() function, and it should work without the missing values error.
Full Corrected Code Snippet (Key Sections)
Here's the fixed version of the critical parts of your code:
# create term document matrix (tdm) tdm <- DocumentTermMatrix(corpus) as.matrix(tdm)[10:20,200:210] # inspect a portion of the tdm # create tf-idf weighted version of term document matrix weightedtdm <- weightTfIdf(tdm) as.matrix(weightedtdm)[10:20,200:210] # inspect same portion of the weighted tdm # find frequent terms: terms that appear in at least "250" documents here, about 25% of the docs findFreqTerms(tdm, 250) # CORRECTLY convert tdm's into data frames tdm <- as.data.frame(as.matrix(tdm)) weightedtdm <- as.data.frame(as.matrix(weightedtdm)) # split back into train and test sets tdmTrain <- tdm[which(merged$train_test == "train"),] weightedTDMtrain <- weightedtdm[which(merged$train_test == "train"),] tdmTest <- tdm[which(merged$train_test == "test"),] weightedTDMtest <- weightedtdm[which(merged$train_test == "test"),] # ... rest of your code stays the same ...
Content of the question comes from Stack Exchange, asked by AndiADL

