tidytext中unnest_token仅token="words"可用,其他参数报错求助
unnest_tokens() with Non-"words" Tokens Hey there! I’ve run into this exact issue before, so let’s break down why this is happening and how to fix it.
What’s Causing the Error?
The token = "words" option works smoothly because it’s handled internally by tidytext, but other token types like "sentences", "ngrams", or "paragraphs" rely on the tokenizers package under the hood. If this dependency is missing, or if you’re running an older version of tidytext, the function can’t properly access your txt column, leading to that frustrating "object not found" error.
Step-by-Step Fixes
Install and Load the
tokenizersDependency
Tidytext doesn’t automatically install tokenizers, so you’ll need to add it manually:install.packages("tokenizers") library(tokenizers)Loading this package ensures the sentence-tokenizing function is available when
unnest_tokens()calls it.Update Tidytext to the Latest Version
Older versions of tidytext had a bug where non-"words" tokens failed to pass the data frame environment correctly. Update with:install.packages("tidytext")Restart your R session after updating to make sure the new version loads properly.
Adjust Your Code (Minor Tweaks)
Try usingtibble()instead ofdata_frame()(they’re nearly identical, buttibble()is the modern preferred syntax in dplyr):library(dplyr) library(janeaustenr) library(tidytext) library(tokenizers) # Create your data frame with tibble() d <- tibble(txt = prideprejudice) # Now try sentence tokenization again d %>% unnest_tokens(sentence, txt, token = "sentences")If that still doesn’t work, you can directly specify the tokenizing function from the tokenizers package to bypass any internal lookup issues:
d %>% unnest_tokens(sentence, txt, token = tokenizers::tokenize_sentences)
Why This Works
When you use token = "sentences", tidytext passes your txt column to tokenizers::tokenize_sentences(). Without the tokenizers package installed, or with an outdated tidytext version, this handoff fails—hence the error about missing txt. By addressing these dependencies and using the latest code, you’ll get the sentence tokenization working just like the word tokenization.
内容的提问来源于stack exchange,提问作者Thomas

