You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用NLTK Naive Bayes实现命名实体识别?附自有语料示例

Using NLTK Naive Bayes for Named Entity Recognition with Your Annotated Corpus

Got it, let's walk through building a NER system with NLTK's Naive Bayes classifier, tailored specifically to your annotated Indonesian-language corpus. Here's a step-by-step implementation you can follow:

Step 1: Parse Your Annotated Corpus

First, we need to convert your text with <ENAMEX> tags into a labeled dataset where each token gets an entity tag (e.g., PERSON, ORGANIZATION, or O for non-entities).

For your sample text:

Sementara itu Pengamat Pasar Modal

return (<ENAMEX TYPE="PERSON">Dandossi Matram</ENAMEX>)
mengatakan, sulit bagi sebuah
return (<ENAMEX TYPE="ORGANIZATION">kantor akuntan publik</ENAMEX>)
(
return (<ENAMEX TYPE="ORGANIZATION">KAP</ENAMEX>)
) untuk dapat menyelesaikan audit perusahaan sebesar
return (<ENAMEX TYPE="ORGANIZATION">Telkom</ENAMEX>)
dalam waktu 3 b...

We'll process it into token-tag pairs like:

[("Sementara", "O"), ("itu", "O"), ("Pengamat", "O"), ("Pasar", "O"), ("Modal", "O"), ("Dandossi", "PERSON"), ("Matram", "PERSON"), ("mengatakan,", "O"), ...]

Here's a Python snippet to parse these tags (tweak the regex if you run into edge cases):

import re
import nltk
from nltk.tokenize import word_tokenize

def parse_enamex_corpus(text):
    # Replace ENAMEX tags with start/end markers to track entity context
    tagged_text = re.sub(r'<ENAMEX TYPE="(\w+)">(.+?)</ENAMEX>', r'__START_\1__ \2 __END__', text)
    tokens = word_tokenize(tagged_text)
    
    token_tags = []
    current_tag = "O"  # Default to non-entity
    
    for token in tokens:
        if token.startswith("__START_"):
            # Extract the entity type from the start marker
            current_tag = token.split("_")[2]
            continue
        if token == "__END__":
            # Reset to non-entity after the end marker
            current_tag = "O"
            continue
        token_tags.append((token, current_tag))
    
    return token_tags

# Test with your sample text
sample_text = 'Sementara itu Pengamat Pasar Modal <ENAMEX TYPE="PERSON">Dandossi Matram</ENAMEX> mengatakan, sulit bagi sebuah <ENAMEX TYPE="ORGANIZATION">kantor akuntan publik</ENAMEX> (<ENAMEX TYPE="ORGANIZATION">KAP</ENAMEX>) untuk dapat menyelesaikan audit perusahaan sebesar <ENAMEX TYPE="ORGANIZATION">Telkom</ENAMEX> dalam waktu 3 b...'
labeled_tokens = parse_enamex_corpus(sample_text)
print(labeled_tokens[:10])

Step 2: Define Feature Functions

Naive Bayes relies on meaningful features to make predictions. For Indonesian NER, we'll use token-level features that capture language-specific patterns:

def extract_features(token, index, tokens):
    features = {
        'token_lower': token.lower(),
        'is_title_case': token.istitle(),
        'is_all_upper': token.isupper(),
        'prefix_2': token[:2].lower(),  # First 2 characters
        'suffix_2': token[-2:].lower(), # Last 2 characters
        'previous_token': tokens[index-1].lower() if index > 0 else '',
        'next_token': tokens[index+1].lower() if index < len(tokens)-1 else '',
        'contains_dash': '-' in token,
    }
    # Optional: Add Indonesian stopword check (uncomment after downloading stopwords)
    # nltk.download('stopwords')
    # indonesian_stopwords = set(nltk.corpus.stopwords.words('indonesian'))
    # features['is_stopword'] = token.lower() in indonesian_stopwords
    return features

# Convert labeled tokens into feature-label pairs for training
def prepare_training_data(labeled_tokens):
    tokens = [t for t, tag in labeled_tokens]
    return [(extract_features(tokens[i], i, tokens), tag) for i, (t, tag) in enumerate(labeled_tokens)]

training_data = prepare_training_data(labeled_tokens)

Step 3: Train the Naive Bayes Classifier

Now we can train the classifier using NLTK's built-in tools, plus evaluate its accuracy:

from nltk.classify import NaiveBayesClassifier

# Split data into train/test sets (use way more data for real-world performance!)
train_size = int(len(training_data) * 0.8)
train_set = training_data[:train_size]
test_set = training_data[train_size:]

# Train the classifier
classifier = NaiveBayesClassifier.train(train_set)

# Calculate and print accuracy
accuracy = nltk.classify.accuracy(classifier, test_set)
print(f"Classifier Accuracy: {accuracy:.2f}")

# Show the most informative features (helps debug what the model learns)
classifier.show_most_informative_features(10)

Step 4: Tag New Unannotated Text

Once trained, you can use the classifier to tag new Indonesian text:

def tag_new_text(text):
    tokens = word_tokenize(text)
    tagged_tokens = []
    for i, token in enumerate(tokens):
        features = extract_features(token, i, tokens)
        tag = classifier.classify(features)
        tagged_tokens.append((token, tag))
    return tagged_tokens

# Example test run
test_text = "Kantor akuntan publik KAP akan melakukan audit Telkom besok"
tagged_result = tag_new_text(test_text)
print(tagged_result)

Key Tips for Better Performance

  • More Data: Your sample is small—Naive Bayes needs sufficient training examples to generalize well. Expand your annotated corpus for better results.
  • Feature Tuning: Add Indonesian-specific features, like checking for common organization prefixes (PT, KAP) or person name suffixes.
  • Entity Merging: The current setup tags individual tokens. To group multi-token entities (like "kantor akuntan publik"), add post-processing logic to merge consecutive tokens with the same non-O tag.

内容的提问来源于stack exchange,提问作者Irfan HP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:45:22