You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何避免俄语短语中店铺专有名词的词形还原

How to Lemmatize Russian Phrases While Preserving Capitalized Store Names

Hey there! Let's fix your lemmatization function so it leaves those capitalized store names untouched. Your original code was on the right track, but it was missing the logic to distinguish proper nouns (your store names) from regular words, and it was also adding all possible lemmas instead of the most accurate one. Here's how to adjust it:

Step-by-Step Solution

First, let's recap the core requirements:

  • Lemmatize regular Russian words (lowercase or non-title-case)
  • Keep capitalized store names (like Авито, Люком) in their original form
  • Skip stopwords and avoid duplicate lemmas

Modified Code

import pymorphy2
from nltk.corpus import stopwords

# Initialize stopwords and morph analyzer
stops = stopwords.words('russian')
morph = pymorphy2.MorphAnalyzer()

def lemmatization_func(x):
    words_from_phrase = x.split()
    unique_lemmas = []
    
    for word in words_from_phrase:
        # Skip stopwords (case-insensitive check)
        if word.lower() in stops:
            continue
        
        # Check if the word is a title-case proper noun (store name)
        if word.istitle():
            processed_word = word  # Keep original store name
        else:
            # Get the most confident lemma for regular words
            parsed_word = morph.parse(word)[0]
            processed_word = parsed_word.normal_form
        
        # Add to list only if not already present (deduplication)
        if processed_word not in unique_lemmas:
            unique_lemmas.append(processed_word)
    
    return ' '.join(unique_lemmas)

# Test with your sample input
x = "купить машину на Авито купить кондиционеры Люком цены на кухни Киргу"
j = lemmatization_func(x)
print(j)

Key Improvements Explained

  1. Proper Noun Detection:
    We use word.istitle() to identify words that start with a capital letter and have the rest in lowercase—this matches exactly the store name format in your example. If you have store names in all caps (like АВИТО), you can add an extra check with word.isupper() to cover that case.

  2. Accurate Lemmatization:
    Instead of looping through all possible parse results (which was adding redundant lemmas), we take the first result from morph.parse(word)—Pymorphy2 returns parses ordered by confidence, so the first one is almost always the correct lemma.

  3. Better Stopword Handling:
    We convert the word to lowercase before checking against the stopwords list, so words like На (capitalized) are still recognized as stopwords and skipped.

  4. Efficient Deduplication:
    We check if the processed word is already in our list before adding it, which keeps the output clean and avoids duplicates without messing up the word order.

Test Output

Running the code with your sample input will produce:

купить машина Авито кондиционер Люком цена кухня Киргу

All regular words are lemmatized correctly, and your store names (Авито, Люком, Киргу) remain unchanged.

Edge Case Considerations

If you encounter regular words that are title-cased (like the start of a sentence), you might need a more nuanced approach—for example, maintaining a predefined list of store names to check against. But for most PPC ad targeting phrases, the title-case check should work perfectly since store names are intentionally capitalized to stand out.

内容的提问来源于stack exchange,提问作者Naglyj.Spamer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:19:21