使用NLTK TaggedCorpusReader提取ANC语料库形容词时遇错误
Hey there! Let's troubleshoot why your adjective extraction code isn't working as expected. Here are a few common issues and fixes based on how the ANC corpus and NLTK's TaggedCorpusReader operate:
1. You're Missing Other Adjective Tag Variants
The ANC uses the Penn Treebank POS tag set, which includes more adjective tags than just JJ. You're likely missing comparative (JJR, e.g., "faster") and superlative (JJS, e.g., "fastest") adjectives with your current exact match condition.
Instead of checking only for JJ, try matching any tag that starts with JJ to capture all adjective forms:
anc_adj = {word.lower() for word, pos in tagged_words if pos.startswith('JJ')}
Or if you want to be explicit about which tags to include:
anc_adj = {word.lower() for word, pos in tagged_words if pos in ('JJ', 'JJR', 'JJS')}
2. The Tags Might Have Extra Suffixes/Prefixes
Some corpora append additional metadata to POS tags (like -NC for non-count categories or -HL for highlighted text). To confirm the actual tag formatting in your corpus, print a sample of tagged words first:
# Inspect the first 20 tagged word pairs to verify tag structure print(tagged_words[:20])
If you see tags like JJ-NC, adjusting your condition to check if the tag starts with JJ (as shown above) will still catch these valid adjective tags.
3. Memory Overload with Large Corpora
The ANC is a substantial corpus, and loading all tagged_words() at once might cause performance slowdowns or memory errors. Process the corpus incrementally using sentences instead to save RAM:
anc_adj = set() # Iterate through each tagged sentence one at a time for sent in anc.tagged_sents(): for word, pos in sent: if pos.startswith('JJ'): anc_adj.add(word.lower())
This approach avoids loading the entire corpus into memory simultaneously, making it more efficient for large datasets.
Give these fixes a try—one of them should resolve your issue!
内容的提问来源于stack exchange,提问作者Jen S

