如何避免俄语短语中店铺专有名词的词形还原
Hey there! Let's fix your lemmatization function so it leaves those capitalized store names untouched. Your original code was on the right track, but it was missing the logic to distinguish proper nouns (your store names) from regular words, and it was also adding all possible lemmas instead of the most accurate one. Here's how to adjust it:
Step-by-Step Solution
First, let's recap the core requirements:
- Lemmatize regular Russian words (lowercase or non-title-case)
- Keep capitalized store names (like Авито, Люком) in their original form
- Skip stopwords and avoid duplicate lemmas
Modified Code
import pymorphy2 from nltk.corpus import stopwords # Initialize stopwords and morph analyzer stops = stopwords.words('russian') morph = pymorphy2.MorphAnalyzer() def lemmatization_func(x): words_from_phrase = x.split() unique_lemmas = [] for word in words_from_phrase: # Skip stopwords (case-insensitive check) if word.lower() in stops: continue # Check if the word is a title-case proper noun (store name) if word.istitle(): processed_word = word # Keep original store name else: # Get the most confident lemma for regular words parsed_word = morph.parse(word)[0] processed_word = parsed_word.normal_form # Add to list only if not already present (deduplication) if processed_word not in unique_lemmas: unique_lemmas.append(processed_word) return ' '.join(unique_lemmas) # Test with your sample input x = "купить машину на Авито купить кондиционеры Люком цены на кухни Киргу" j = lemmatization_func(x) print(j)
Key Improvements Explained
Proper Noun Detection:
We useword.istitle()to identify words that start with a capital letter and have the rest in lowercase—this matches exactly the store name format in your example. If you have store names in all caps (likeАВИТО), you can add an extra check withword.isupper()to cover that case.Accurate Lemmatization:
Instead of looping through all possible parse results (which was adding redundant lemmas), we take the first result frommorph.parse(word)—Pymorphy2 returns parses ordered by confidence, so the first one is almost always the correct lemma.Better Stopword Handling:
We convert the word to lowercase before checking against the stopwords list, so words likeНа(capitalized) are still recognized as stopwords and skipped.Efficient Deduplication:
We check if the processed word is already in our list before adding it, which keeps the output clean and avoids duplicates without messing up the word order.
Test Output
Running the code with your sample input will produce:
купить машина Авито кондиционер Люком цена кухня Киргу
All regular words are lemmatized correctly, and your store names (Авито, Люком, Киргу) remain unchanged.
Edge Case Considerations
If you encounter regular words that are title-cased (like the start of a sentence), you might need a more nuanced approach—for example, maintaining a predefined list of store names to check against. But for most PPC ad targeting phrases, the title-case check should work perfectly since store names are intentionally capitalized to stand out.
内容的提问来源于stack exchange,提问作者Naglyj.Spamer

