R中textstem包词形还原异常:"dive"被转为"diva"求助
It looks like the issue comes from a wrong mapping in the default lexicon::hash_lemmas dictionary—somehow the lemma entry for the base word "dive" is set to "diva" instead of "dive" itself. Here's how you can fix this for both your test case and your large 80k+ comment dataset:
Step 1: Confirm the Problem
First, let's check the faulty mapping in the original dictionary to be sure:
library(lexicon) hash_lemmas[["dive"]] # This will return "diva", which is the root cause
Step 2: Build a Corrected Dictionary
We'll create a small custom correction to override the wrong entry, then merge it with the original dictionary. Custom entries will take priority over default ones:
library(textstem) library(lexicon) # Define your targeted lemma fixes custom_corrections <- c( "dive" = "dive" # Correct the base word's lemma mapping ) # Merge with the original dictionary (custom entries overwrite defaults) fixed_lemma_dict <- c(hash_lemmas, custom_corrections)
Step 3: Test the Fixed Lemmatization
Run your test case again with the corrected dictionary to verify it works:
words <- c("dived", "diving", "dive") lemmatize_strings(words, dictionary = fixed_lemma_dict) # Expected output: [1] "dive" "dive" "dive"
Step 4: Apply to Your Large Dataset
For your 80k+ comments, use the fixed dictionary exactly as you would the original one—no extra heavy lifting needed:
# Assuming your comments are stored in a vector named `comments` lemmatized_comments <- lemmatize_strings(comments, dictionary = fixed_lemma_dict)
If you run into other incorrect lemmatizations later, just add more entries to custom_corrections. This approach is lightweight and flexible, making it ideal for processing large text datasets efficiently.
内容的提问来源于stack exchange,提问作者Kenpy

