如何重新训练Spacy现有NER模型识别多国货币实体
Got it, let's tackle this problem step by step. You're totally right that the basic fixed-string entity tutorial doesn't cut it here—currency mentions are tricky because they show up in so many formats (paired with numbers, abbreviations, full names, different placements). With your 1000 samples, we can fine-tune spaCy's en_core_web_sm model to reliably recognize all those currency variants you listed. Here's how to do it:
SpaCy's NER training expects data in a specific structure: a list of tuples where each tuple is (text, {"entities": [(start_char, end_char, label)]}).
The key here is to consistently label every currency-related token/phrase you care about. For example:
- For
"Rs. 10,000", label"Rs."asCURRENCY:(0, 3, "CURRENCY") - For
"1 million euros", label"euros"asCURRENCY:(10, 15, "CURRENCY") - For
"INR 1 thousand", label"INR"asCURRENCY:(0, 3, "CURRENCY")
If you haven't already, use a tool like Prodigy (spaCy's official annotation tool) to speed up labeling—manually formatting 1000 samples is tedious. If Prodigy isn't an option, you can write a quick script to convert your existing data into this structure.
The en_core_web_sm model already has a CURRENCY entity label built-in—it just doesn't cover all the variants you need. We'll do incremental training to teach it the new patterns while preserving its existing NER skills.
Here's a practical, commented script you can adapt:
import spacy from spacy.training import Example from spacy.util import minibatch, compounding import random # Load the base model nlp = spacy.load("en_core_web_sm") # Grab the NER component (this is what we'll train) ner = nlp.get_pipe("ner") # Double-check the CURRENCY label exists (it should in en_core_web_sm, but just in case) ner.add_label("CURRENCY") # Replace this with your full 1000-sample training data training_data = [ ("Rs. 10,000", {"entities": [(0, 3, "CURRENCY")]}), ("1 million euros", {"entities": [(10, 15, "CURRENCY")]}), ("INR 1 thousand", {"entities": [(0, 3, "CURRENCY")]}), ("eu 500", {"entities": [(0, 2, "CURRENCY")]}), ("200 rupees", {"entities": [(4, 10, "CURRENCY")]}), # Add all your remaining samples here ] # Split into training (80%) and validation (20%) sets to avoid overfitting random.shuffle(training_data) train_set = training_data[:800] dev_set = training_data[800:] # Disable other pipeline components (parser, tagger) during training to save resources other_pipes = [pipe for pipe in nlp.pipe_names if pipe != "ner"] with nlp.disable_pipes(*other_pipes): # Resume training from the existing model's weights optimizer = nlp.resume_training() optimizer.alpha = 0.001 # Lower learning rate for incremental training optimizer.max_grad_norm = 1.0 # Train for 10 iterations (adjust based on loss trends) n_iter = 10 for itn in range(n_iter): random.shuffle(train_set) losses = {} # Use dynamic batch sizes to optimize training batches = minibatch(train_set, size=compounding(4.0, 32.0, 1.001)) for batch in batches: texts, annotations = zip(*batch) # Convert texts and annotations to spaCy Example objects examples = [Example.from_dict(nlp.make_doc(text), annot) for text, annot in zip(texts, annotations)] # Update the model nlp.update( examples, sgd=optimizer, drop=0.2, # Dropout to prevent overfitting losses=losses, ) print(f"Iteration {itn+1} | Loss: {losses['ner']:.4f}") # Evaluate the model on the validation set from spacy.scorer import Scorer scorer = Scorer() eval_examples = [] for text, annot in dev_set: doc = nlp.make_doc(text) example = Example.from_dict(doc, annot) example.predicted = nlp(text) eval_examples.append(example) scores = scorer.score(eval_examples) print("\nValidation Set Performance for CURRENCY:") print(f"Precision: {scores['ents_per_type']['CURRENCY']['p']:.4f}") print(f"Recall: {scores['ents_per_type']['CURRENCY']['r']:.4f}") print(f"F1 Score: {scores['ents_per_type']['CURRENCY']['f']:.4f}") # Save your fine-tuned model for later use nlp.to_disk("./en_core_web_sm_currency")
Even with 1000 samples, there might be rare currency variants the model misses. SpaCy's Matcher lets you add hard-coded rules to catch these, which you can combine with the model's predictions. Here's how:
from spacy.matcher import Matcher # Initialize the matcher with the model's vocab matcher = Matcher(nlp.vocab) # Define patterns for all your currency variants currency_patterns = [ [{"LOWER": {"IN": ["rs", "rs.", "inr", "eur", "eu", "euro", "euros", "rupee", "rupees"]}}], [{"TEXT": {"IN": ["€", "₹"]}}] # Add currency symbols if needed ] # Add the patterns to the matcher for pattern in currency_patterns: matcher.add("CURRENCY_MATCH", [pattern]) # Create a custom pipeline component to add entities from matcher results def add_currency_entities(doc): matches = matcher(doc) for match_id, start, end in matches: span = doc[start:end] # Only add the entity if it's not already labeled by the NER model if not span.ents: span_ent = span.as_doc()[0].set_ent_type_("CURRENCY") doc.ents = list(doc.ents) + [span_ent] return doc # Add this component after the NER pipe in your model nlp.add_pipe("add_currency_entities", after="ner")
After your first training run, test the model on unseen text to find gaps. Did it miss "eu" in a weird context? Did it fail to label "rupees" when it's at the start of a sentence? Add those missed cases to your training data and re-run the training loop. With 1000 samples, a few rounds of iteration will make the model really robust.
内容的提问来源于stack exchange,提问作者Lau Kumra

