如何使用NLTK Naive Bayes实现命名实体识别?附自有语料示例
Got it, let's walk through building a NER system with NLTK's Naive Bayes classifier, tailored specifically to your annotated Indonesian-language corpus. Here's a step-by-step implementation you can follow:
Step 1: Parse Your Annotated Corpus
First, we need to convert your text with <ENAMEX> tags into a labeled dataset where each token gets an entity tag (e.g., PERSON, ORGANIZATION, or O for non-entities).
For your sample text:
Sementara itu Pengamat Pasar Modal
mengatakan, sulit bagi sebuahreturn (<ENAMEX TYPE="PERSON">Dandossi Matram</ENAMEX>)(return (<ENAMEX TYPE="ORGANIZATION">kantor akuntan publik</ENAMEX>)) untuk dapat menyelesaikan audit perusahaan sebesarreturn (<ENAMEX TYPE="ORGANIZATION">KAP</ENAMEX>)dalam waktu 3 b...return (<ENAMEX TYPE="ORGANIZATION">Telkom</ENAMEX>)
We'll process it into token-tag pairs like:
[("Sementara", "O"), ("itu", "O"), ("Pengamat", "O"), ("Pasar", "O"), ("Modal", "O"), ("Dandossi", "PERSON"), ("Matram", "PERSON"), ("mengatakan,", "O"), ...]
Here's a Python snippet to parse these tags (tweak the regex if you run into edge cases):
import re import nltk from nltk.tokenize import word_tokenize def parse_enamex_corpus(text): # Replace ENAMEX tags with start/end markers to track entity context tagged_text = re.sub(r'<ENAMEX TYPE="(\w+)">(.+?)</ENAMEX>', r'__START_\1__ \2 __END__', text) tokens = word_tokenize(tagged_text) token_tags = [] current_tag = "O" # Default to non-entity for token in tokens: if token.startswith("__START_"): # Extract the entity type from the start marker current_tag = token.split("_")[2] continue if token == "__END__": # Reset to non-entity after the end marker current_tag = "O" continue token_tags.append((token, current_tag)) return token_tags # Test with your sample text sample_text = 'Sementara itu Pengamat Pasar Modal <ENAMEX TYPE="PERSON">Dandossi Matram</ENAMEX> mengatakan, sulit bagi sebuah <ENAMEX TYPE="ORGANIZATION">kantor akuntan publik</ENAMEX> (<ENAMEX TYPE="ORGANIZATION">KAP</ENAMEX>) untuk dapat menyelesaikan audit perusahaan sebesar <ENAMEX TYPE="ORGANIZATION">Telkom</ENAMEX> dalam waktu 3 b...' labeled_tokens = parse_enamex_corpus(sample_text) print(labeled_tokens[:10])
Step 2: Define Feature Functions
Naive Bayes relies on meaningful features to make predictions. For Indonesian NER, we'll use token-level features that capture language-specific patterns:
def extract_features(token, index, tokens): features = { 'token_lower': token.lower(), 'is_title_case': token.istitle(), 'is_all_upper': token.isupper(), 'prefix_2': token[:2].lower(), # First 2 characters 'suffix_2': token[-2:].lower(), # Last 2 characters 'previous_token': tokens[index-1].lower() if index > 0 else '', 'next_token': tokens[index+1].lower() if index < len(tokens)-1 else '', 'contains_dash': '-' in token, } # Optional: Add Indonesian stopword check (uncomment after downloading stopwords) # nltk.download('stopwords') # indonesian_stopwords = set(nltk.corpus.stopwords.words('indonesian')) # features['is_stopword'] = token.lower() in indonesian_stopwords return features # Convert labeled tokens into feature-label pairs for training def prepare_training_data(labeled_tokens): tokens = [t for t, tag in labeled_tokens] return [(extract_features(tokens[i], i, tokens), tag) for i, (t, tag) in enumerate(labeled_tokens)] training_data = prepare_training_data(labeled_tokens)
Step 3: Train the Naive Bayes Classifier
Now we can train the classifier using NLTK's built-in tools, plus evaluate its accuracy:
from nltk.classify import NaiveBayesClassifier # Split data into train/test sets (use way more data for real-world performance!) train_size = int(len(training_data) * 0.8) train_set = training_data[:train_size] test_set = training_data[train_size:] # Train the classifier classifier = NaiveBayesClassifier.train(train_set) # Calculate and print accuracy accuracy = nltk.classify.accuracy(classifier, test_set) print(f"Classifier Accuracy: {accuracy:.2f}") # Show the most informative features (helps debug what the model learns) classifier.show_most_informative_features(10)
Step 4: Tag New Unannotated Text
Once trained, you can use the classifier to tag new Indonesian text:
def tag_new_text(text): tokens = word_tokenize(text) tagged_tokens = [] for i, token in enumerate(tokens): features = extract_features(token, i, tokens) tag = classifier.classify(features) tagged_tokens.append((token, tag)) return tagged_tokens # Example test run test_text = "Kantor akuntan publik KAP akan melakukan audit Telkom besok" tagged_result = tag_new_text(test_text) print(tagged_result)
Key Tips for Better Performance
- More Data: Your sample is small—Naive Bayes needs sufficient training examples to generalize well. Expand your annotated corpus for better results.
- Feature Tuning: Add Indonesian-specific features, like checking for common organization prefixes (
PT,KAP) or person name suffixes. - Entity Merging: The current setup tags individual tokens. To group multi-token entities (like "kantor akuntan publik"), add post-processing logic to merge consecutive tokens with the same non-
Otag.
内容的提问来源于stack exchange,提问作者Irfan HP

