R语言文本挖掘报错:二元运算符的非数值参数问题求助
non-numeric argument to binary operator Error in Your Text Analysis Workflow Hey there! Let's break down that frustrating error you're hitting. I know it's super annoying when you're following a book step-by-step, have already checked Google, the book's errata, and Stack Overflow, and still hit a roadblock. Let's figure this out together.
First, let's clarify what that error means: it pops up when you try to use a binary operator (like +, -, *, /, or even %>% in some edge cases) with at least one non-numeric value. For example, if you try to add a character string to a number, or multiply a list by a vector—R doesn't know how to handle that, so it throws this error.
Context from Your Workflow
You mentioned these lines run fine:
original_books <- austen_books() %>% group_by(book) %>% mutate(linenumber = row_number(), chapter = cumsum(str_detect(text, regex("^chapter [\\divxlc]", ignore_case = TRUE)))) %>% ungroup() tidy_books <- original_books %>% unnest_tokens(word, text)
So the issue is definitely in the subsequent code you ran after creating tidy_books. Let's walk through the most likely culprits and how to debug them:
Step 1: Pinpoint the Exact Line Causing the Error
First, split your follow-up code into individual lines and run them one by one. This will tell you exactly which operation is triggering the error. For example, if you had a block like:
tidy_books <- tidy_books %>% anti_join(stop_words) %>% count(word) %>% mutate(ratio = n / total_words)
Run each pipe segment separately to see where it breaks.
Step 2: Check Variable Types
Once you've found the problematic line, use the class() function to check the type of every object involved in the binary operation. For example:
- If you're running
ratio = n / total_words, runclass(n)andclass(total_words)—one of them is probably a character vector, list, orNULLinstead of a numeric value. - If you're using
cumsum()orsum()somewhere, make sure the input to those functions is numeric (not character).
Step 3: Common Pitfalls in This Text Analysis Workflow
Given you're working with the tidytext framework (from the looks of austen_books() and unnest_tokens()), here are the most common places this error pops up:
- Miscalculating totals: If you're trying to divide by a total word count, make sure you've actually calculated that total as a numeric value (e.g.,
total_words <- sum(tidy_books_counted$n)instead of accidentally storing a string like "total words"). - Stop word join issues: If you ran
anti_join(stop_words), double-check that thewordcolumn in bothtidy_booksandstop_wordsis the same type (both character—this is usually fine, but sometimes typos can break things). - TF-IDF missteps: When using
bind_tf_idf(), ensure you're passing the correct columns (word, document, count) in the right order. If you mix up columns, you might end up with non-numeric values intforidf, leading to errors when multiplying them. - Accidental character conversion: If you modified the
tidy_booksdata frame with a function likemutate()that converted a numeric column to character (e.g., adding a string suffix by mistake), that would break any subsequent numeric operations.
Step 4: Double-Check Book Code Details
Even with checking errata, it's easy to miss a small detail:
- Did you use the exact variable names the book uses? (e.g.,
linenumbervs.line_number,chaptervs.chapters) - Are you using the same package versions as the book? Sometimes updates to tidytext, dplyr, or stringr can change behavior slightly (you can check package versions with
packageVersion("tidytext")).
If you can share the exact lines of code that triggered the error, we can narrow this down even further. But with these steps, you should be able to track down the source of the non-numeric argument issue.
内容的提问来源于stack exchange,提问作者heeby89

