如何修改NLTK分词代码以按词性分类输出结果?
Solution for Categorizing Nouns and Verbs with NLTK
Step-by-Step Implementation
- First, ensure you've downloaded required NLTK resources (run once):
import nltk nltk.download('punkt') nltk.download('averaged_perceptron_tagger') - Use NLTK's
word_tokenizeto split text into words, thenpos_tagto get part-of-speech tags. - Filter tags to separate nouns and verbs:
- Nouns correspond to tags starting with
NN(e.g., NN, NNS, NNP, NNPS) - Verbs correspond to tags starting with
VB(e.g., VB, VBD, VBG, VBN, VBP, VBZ)
- Nouns correspond to tags starting with
- Format results into the requested list syntax.
Complete Code Example
import nltk from nltk.tokenize import word_tokenize from nltk.tag import pos_tag # Download required resources (run once) nltk.download('punkt') nltk.download('averaged_perceptron_tagger') def categorize_pos(text): # Tokenize and tag input text tokens = word_tokenize(text) tagged_tokens = pos_tag(tokens) # Separate nouns and verbs based on POS tags nouns = [word for word, tag in tagged_tokens if tag.startswith('NN')] verbs = [word for word, tag in tagged_tokens if tag.startswith('VB')] # Output in the requested format print(f'nouns = {nouns}') print(f'verbs = {verbs}') # Test with sample text sample_text = "Natural language processing is fascinating" categorize_pos(sample_text)
Example Output
nouns = ['Natural', 'language', 'processing'] verbs = ['is', 'fascinating']
Customization Notes
- Adjust tag filters if you need precise control (e.g., exclude proper nouns using
tag in ['NN', 'NNS']instead ofstartswith('NN')). - The output directly uses valid Python list syntax as requested.
内容的提问来源于stack exchange,提问作者Ali
相关产品推荐
相关产品推荐

