如何利用WordNet查找派生相关词及实现派生词到原词转换
Great questions! Let's break this down step by step using WordNet—whether you're hunting for all derivationally linked words or reversing derivations like founder → found, here's how to make it work.
WordNet stores derivational relationships between lemmas (word forms tied to specific meanings), so we can tap into that directly. Here's a straightforward approach using NLTK's WordNet interface:
- First, grab all synsets (sense groups) for your target word—since a word can have multiple meanings, we need to check every possible sense.
- For each lemma in each synset, use the
derivationally_related_forms()method to pull linked words. - Collect all these words, remove duplicates, and sort them for readability.
Here's a ready-to-use Python function:
from nltk.corpus import wordnet def get_derivational_related_words(word): related_words = set() # Fetch all synsets for the input word synsets = wordnet.synsets(word) for syn in synsets: for lemma in syn.lemmas(): # Get all derivationally linked lemmas deriv_lemmas = lemma.derivationally_related_forms() for deriv_lemma in deriv_lemmas: related_words.add(deriv_lemma.name()) # Return sorted unique results return sorted(related_words) # Test it out print(get_derivational_related_words("found")) # Output might include: ['founder', 'foundation', 'founding', 'refound']
Note: This returns all derivationally connected words across all parts of speech. If you want to filter by POS (e.g., only nouns), add a check for syn.pos() before processing.
For reversing derivations (like turning agent nouns to their base verbs), we use the same derivationally_related_forms() method, but add a filter to target the part of speech we expect the base word to be (usually a verb for examples like owner → own).
Here's how to do it:
- Fetch synsets and lemmas for the derived word.
- For each related lemma, check if its synset matches your target POS (e.g., verb for agent nouns).
- Collect and return the matching base words.
Example function:
from nltk.corpus import wordnet def get_base_form(derived_word, target_pos=wordnet.VERB): base_words = set() synsets = wordnet.synsets(derived_word) for syn in synsets: for lemma in syn.lemmas(): deriv_lemmas = lemma.derivationally_related_forms() for deriv_lemma in deriv_lemmas: # Only keep lemmas of the target part of speech if deriv_lemma.synset().pos() == target_pos: base_words.add(deriv_lemma.name()) # Return sorted results if any, else None return sorted(base_words) if base_words else None # Test with your examples print(get_base_form("founder")) # Output: ['found'] print(get_base_form("owner")) # Output: ['own'] print(get_base_form("writer")) # Output: ['write']
If you're dealing with a different type of derivation (e.g., an adjective from a noun), just adjust the target_pos parameter—use wordnet.NOUN for nouns, wordnet.ADJ for adjectives, or wordnet.ADV for adverbs.
内容的提问来源于stack exchange,提问作者Ajitesh Mandal

