如何使用NLTK合并词性标注后的连续专有名词?
Got it, let's fix that issue where you're only extracting single proper nouns instead of merging consecutive ones like "Ronald McDonald" or "Central Park". Here's a simple, effective approach tailored to your use case:
Step-by-Step Solution
First, let's start with your existing code to get the tagged tokens, then add logic to group consecutive NNP tags into full phrases:
import nltk # Download required resources if you haven't already nltk.download('punkt') nltk.download('averaged_perceptron_tagger') sentence = 'Ronald McDonald exercised at the Central Park on Monday.' tokens = nltk.word_tokenize(sentence) tagged = nltk.pos_tag(tokens) # Initialize lists to track and store merged proper nouns merged_proper_nouns = [] current_consecutive_nnp = [] # Iterate through each tagged word pair for word, tag in tagged: # If the word is a proper noun, add it to our temporary list if tag == 'NNP': current_consecutive_nnp.append(word) else: # If we were collecting proper nouns, merge and save them if current_consecutive_nnp: merged_phrase = ' '.join(current_consecutive_nnp) merged_proper_nouns.append(merged_phrase) current_consecutive_nnp = [] # Reset for next group # Don't forget to catch any remaining proper nouns at the end of the sentence if current_consecutive_nnp: merged_proper_nouns.append(' '.join(current_consecutive_nnp)) print(merged_proper_nouns) # Output: ['Ronald McDonald', 'Central Park', 'Monday']
How This Works
- Temporary Tracking: We use
current_consecutive_nnpto hold words that are consecutive proper nouns as we iterate through the tagged list. - Merge & Save: When we hit a non-proper noun, we check if we've been collecting NNPs. If so, we join them into a single phrase and add it to our final list, then reset the temporary list.
- Final Check: After the loop ends, we make sure to add any remaining NNPs (in case the sentence ends with a proper noun, like "Monday" in your example).
Bonus: Handle Plural Proper Nouns
If you also want to include plural proper nouns tagged as NNPS, just adjust the condition to check if the tag starts with "NNP":
if tag.startswith('NNP'): current_consecutive_nnp.append(word)
This will catch both singular (NNP) and plural (NNPS) proper nouns, so phrases like "The Smiths" will be merged correctly too.
内容的提问来源于stack exchange,提问作者user9092346
相关产品推荐
相关产品推荐

