You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何重新训练Spacy现有NER模型识别多国货币实体

Got it, let's tackle this problem step by step. You're totally right that the basic fixed-string entity tutorial doesn't cut it here—currency mentions are tricky because they show up in so many formats (paired with numbers, abbreviations, full names, different placements). With your 1000 samples, we can fine-tune spaCy's en_core_web_sm model to reliably recognize all those currency variants you listed. Here's how to do it:

1. First, Get Your Training Data in SpaCy's Format

SpaCy's NER training expects data in a specific structure: a list of tuples where each tuple is (text, {"entities": [(start_char, end_char, label)]}).

The key here is to consistently label every currency-related token/phrase you care about. For example:

  • For "Rs. 10,000", label "Rs." as CURRENCY: (0, 3, "CURRENCY")
  • For "1 million euros", label "euros" as CURRENCY: (10, 15, "CURRENCY")
  • For "INR 1 thousand", label "INR" as CURRENCY: (0, 3, "CURRENCY")

If you haven't already, use a tool like Prodigy (spaCy's official annotation tool) to speed up labeling—manually formatting 1000 samples is tedious. If Prodigy isn't an option, you can write a quick script to convert your existing data into this structure.

2. Fine-Tune the Pre-Trained Model (Don't Train From Scratch)

The en_core_web_sm model already has a CURRENCY entity label built-in—it just doesn't cover all the variants you need. We'll do incremental training to teach it the new patterns while preserving its existing NER skills.

3. Write the Training Script

Here's a practical, commented script you can adapt:

import spacy
from spacy.training import Example
from spacy.util import minibatch, compounding
import random

# Load the base model
nlp = spacy.load("en_core_web_sm")

# Grab the NER component (this is what we'll train)
ner = nlp.get_pipe("ner")

# Double-check the CURRENCY label exists (it should in en_core_web_sm, but just in case)
ner.add_label("CURRENCY")

# Replace this with your full 1000-sample training data
training_data = [
    ("Rs. 10,000", {"entities": [(0, 3, "CURRENCY")]}),
    ("1 million euros", {"entities": [(10, 15, "CURRENCY")]}),
    ("INR 1 thousand", {"entities": [(0, 3, "CURRENCY")]}),
    ("eu 500", {"entities": [(0, 2, "CURRENCY")]}),
    ("200 rupees", {"entities": [(4, 10, "CURRENCY")]}),
    # Add all your remaining samples here
]

# Split into training (80%) and validation (20%) sets to avoid overfitting
random.shuffle(training_data)
train_set = training_data[:800]
dev_set = training_data[800:]

# Disable other pipeline components (parser, tagger) during training to save resources
other_pipes = [pipe for pipe in nlp.pipe_names if pipe != "ner"]
with nlp.disable_pipes(*other_pipes):
    # Resume training from the existing model's weights
    optimizer = nlp.resume_training()
    optimizer.alpha = 0.001  # Lower learning rate for incremental training
    optimizer.max_grad_norm = 1.0

    # Train for 10 iterations (adjust based on loss trends)
    n_iter = 10
    for itn in range(n_iter):
        random.shuffle(train_set)
        losses = {}
        # Use dynamic batch sizes to optimize training
        batches = minibatch(train_set, size=compounding(4.0, 32.0, 1.001))
        for batch in batches:
            texts, annotations = zip(*batch)
            # Convert texts and annotations to spaCy Example objects
            examples = [Example.from_dict(nlp.make_doc(text), annot) for text, annot in zip(texts, annotations)]
            # Update the model
            nlp.update(
                examples,
                sgd=optimizer,
                drop=0.2,  # Dropout to prevent overfitting
                losses=losses,
            )
        print(f"Iteration {itn+1} | Loss: {losses['ner']:.4f}")

# Evaluate the model on the validation set
from spacy.scorer import Scorer
scorer = Scorer()
eval_examples = []
for text, annot in dev_set:
    doc = nlp.make_doc(text)
    example = Example.from_dict(doc, annot)
    example.predicted = nlp(text)
    eval_examples.append(example)

scores = scorer.score(eval_examples)
print("\nValidation Set Performance for CURRENCY:")
print(f"Precision: {scores['ents_per_type']['CURRENCY']['p']:.4f}")
print(f"Recall: {scores['ents_per_type']['CURRENCY']['r']:.4f}")
print(f"F1 Score: {scores['ents_per_type']['CURRENCY']['f']:.4f}")

# Save your fine-tuned model for later use
nlp.to_disk("./en_core_web_sm_currency")
4. Add Rule-Based Matching to Cover Edge Cases

Even with 1000 samples, there might be rare currency variants the model misses. SpaCy's Matcher lets you add hard-coded rules to catch these, which you can combine with the model's predictions. Here's how:

from spacy.matcher import Matcher

# Initialize the matcher with the model's vocab
matcher = Matcher(nlp.vocab)

# Define patterns for all your currency variants
currency_patterns = [
    [{"LOWER": {"IN": ["rs", "rs.", "inr", "eur", "eu", "euro", "euros", "rupee", "rupees"]}}],
    [{"TEXT": {"IN": ["€", "₹"]}}]  # Add currency symbols if needed
]

# Add the patterns to the matcher
for pattern in currency_patterns:
    matcher.add("CURRENCY_MATCH", [pattern])

# Create a custom pipeline component to add entities from matcher results
def add_currency_entities(doc):
    matches = matcher(doc)
    for match_id, start, end in matches:
        span = doc[start:end]
        # Only add the entity if it's not already labeled by the NER model
        if not span.ents:
            span_ent = span.as_doc()[0].set_ent_type_("CURRENCY")
            doc.ents = list(doc.ents) + [span_ent]
    return doc

# Add this component after the NER pipe in your model
nlp.add_pipe("add_currency_entities", after="ner")
5. Iterate and Refine

After your first training run, test the model on unseen text to find gaps. Did it miss "eu" in a weird context? Did it fail to label "rupees" when it's at the start of a sentence? Add those missed cases to your training data and re-run the training loop. With 1000 samples, a few rounds of iteration will make the model really robust.


内容的提问来源于stack exchange,提问作者Lau Kumra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:32:49