R语言调用textEmbed函数报错求助:下标越界与未初始化列`words`问题
textEmbed() Error: "subscript out of bounds" & "Unknown or uninitialised column: words" Hey there, let's break down why your textEmbed() call stopped working suddenly and fix it step by step. First, let's recap your scenario:
- You had this code running perfectly before:
tmp1 <- textEmbed(x = 'sofa help', model = 'roberta-base', layers = 11) tmp1$x - Now it throws:
Error in x[[1]] : subscript out of bounds
In addition: Warning message: Unknown or uninitialised column:words.
Your environment setup includes all the required packages, so let's dig into the likely culprits:
1. Fix the udpipe Model & Text Preprocessing Issue
The Unknown or uninitialised column: 'words' warning is the key here—it means the tokenization step (handled by udpipe) isn't producing the expected output. Let's verify and fix this:
- First, test if udpipe is working correctly on its own:
test_text <- "sofa help" # Run udpipe annotation directly udpipe_output <- udpipe_annotate(ud_eng, x = test_text) udpipe_df <- as.data.frame(udpipe_output) # Check if 'words' column exists print(colnames(udpipe_df)) - If
wordsis missing, your model path is probably wrong. Re-download and reload the udpipe model to be safe:# Download the model again (uncommented) udpipe_download_model(language = "english") # Load using the exact filename (avoid here::here if path is ambiguous) ud_eng <- udpipe_load_model("english-ewt-ud-2.5-191206.udpipe")
2. Reset Your Reticulate Python Environment
Since textEmbed() relies on reticulate to call PyTorch/transformers, a corrupted Python session or dependency conflict could be causing the issue.
- Restart your R session first (this is crucial—don't skip it). Then re-run your setup code cleanly:
require(quanteda) require(quanteda.textstats) require(udpipe) require(reticulate) require(text) # Verify your Python environment is correctly loaded py_config() # Re-import the required Python libraries reticulate::import('torch') reticulate::import('transformers') - If your Python environment looks off, explicitly specify your virtual environment:
# Replace with your actual virtual environment name use_virtualenv("your_virtualenv_name", required = TRUE)
3. Check for text Package Updates or Cache Corruption
Sometimes package updates or cached model files can break functionality:
- Check your
textpackage version and update it if needed:packageVersion("text") # Update to the latest stable version install.packages("text") - If the issue persists, clear the
textpackage's cached models (this will force it to re-download the RoBERTa model):# Locate the cache directory text_cache <- system.file("cache", package = "text") # Delete the cache (only do this if you're okay with re-downloading models) unlink(text_cache, recursive = TRUE)
4. Simplify Testing to Narrow Down the Problem
Start with the most basic textEmbed() call to rule out parameter issues:
# Test with the default model first tmp_test <- textEmbed(x = "sofa help") tmp_test$x
If that works, try your original model again—maybe the layers parameter was causing a niche issue (roberta-base has 12 layers, so try layers = 12 just in case):
tmp1 <- textEmbed(x = 'sofa help', model = 'roberta-base', layers = 12) tmp1$x
内容的提问来源于stack exchange,提问作者vinita vader

