You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Spacy NER实体标签过滤训练数据:移除无实体句子的实现问题

Solution: Filter out entity-free sentences and adjust entity indices for spaCy NER training

Got it, let's tackle this problem step by step. The core challenge here is twofold: splitting your text into sentences based on periods and newlines, then filtering out sentences that don't contain any entities—and adjusting the entity character indices since they're tied to the original full text, not the split sentences.

Here's a complete implementation that builds on your existing code:

from tqdm import tqdm
import spacy
from spacy.tokens import DocBin

def split_text_into_sentences_with_indices(text):
    """Split text into sentences using . and \n as delimiters, return each sentence with its original start/end indices."""
    sentences = []
    current_start = 0
    i = 0
    while i < len(text):
        # Check for sentence delimiters
        if text[i] in ('.', '\n'):
            # Extract the sentence (trim whitespace to avoid empty sentences)
            sentence_text = text[current_start:i].strip()
            if sentence_text:
                sentences.append((sentence_text, current_start, i))
            # Move past the delimiter
            current_start = i + 1
        i += 1
    # Add the last sentence if there's any remaining text
    final_sentence = text[current_start:].strip()
    if final_sentence:
        sentences.append((final_sentence, current_start, len(text)))
    return sentences

nlp = spacy.blank("en")
db = DocBin()
train_data = [('Christmas Perot 2021 TSO\nSkip to Main Content HOME CONCERTS EVENTS ABOUT STAFF EDUCATION SUPPORT US More Use tab to navigate through the menu items. BUY TICKETS SUNDAY, DECEMBER 12, 2021 I PEROT THEATRE I 4:00 PM\nPOPS I Christmas at The Perot\nCLICK HERE to purchase tickets, or contact the Texarkana Symphony Orchestra at 870.773.3401\nA Texarkana Tradition Join the TSO, the Texarkana Jazz Orchestra, and the TSO Chamber Singers, for this holiday concert for the whole family.\nDon’t miss seeing the winner of TSO’s 11th Annual Celebrity Conductor Competition\nBack to Events 2019 Texarkana Symphony Orchestra', {'entities': [(375, 399, 'organization'), (290, 318, 'organization'), (220, 242, 'production_name'), (169, 186, 'performance_date'), (189, 202, 'auditorium'), (205, 212, 'performance_starttime'), (409, 428, 'organization')]})]

for text, annot in tqdm(train_data):
    # Split text into sentences with their original positions
    sentences_with_indices = split_text_into_sentences_with_indices(text)
    
    for sent_text, sent_start, sent_end in sentences_with_indices:
        # Check if any entity falls within this sentence's bounds
        sent_entities = []
        for ent_start, ent_end, ent_label in annot["entities"]:
            # Entity is fully contained within the sentence
            if sent_start <= ent_start < ent_end <= sent_end:
                # Adjust entity indices to be relative to the sentence start
                rel_start = ent_start - sent_start
                rel_end = ent_end - sent_start
                sent_entities.append((rel_start, rel_end, ent_label))
        
        # Only process sentences that have at least one entity
        if sent_entities:
            doc = nlp.make_doc(sent_text)
            ents = []
            for start, end, label in sent_entities:
                span = doc.char_span(start, end, label=label, alignment_mode="contract")
                if span is None:
                    print(f"Skipping entity ({start}, {end}, {label}) in sentence: {sent_text}")
                else:
                    ents.append(span)
            doc.ents = ents
            db.add(doc)

# Save the processed DocBin if needed
# db.to_disk("./filtered_ner_data.spacy")

Let's break down what this does:

  1. Sentence Splitting with Indices:

    • The split_text_into_sentences_with_indices function splits your text using . and \n as delimiters, and keeps track of each sentence's starting and ending character positions in the original full text. This is crucial for matching entities to sentences.
    • It also trims whitespace to avoid empty sentences from consecutive delimiters.
  2. Filtering & Index Adjustment:

    • For each sentence, we check if any of the original entities are fully contained within the sentence's bounds.
    • If an entity belongs to the sentence, we subtract the sentence's starting index from the entity's original start/end to get relative indices for the sentence.
    • We only create a Doc object for sentences that have at least one entity, skipping the ones with no entities entirely.
  3. Integration with spaCy:

    • The rest of the code follows your original workflow: creating Doc objects, adding adjusted entities, and populating the DocBin.

Testing with your sample data:

The first line Christmas Perot 2021 TSO has no entities, so it gets skipped entirely. Sentences containing entities like BUY TICKETS SUNDAY, DECEMBER 12, 2021 I PEROT THEATRE I 4:00 PM will be retained, with their entities adjusted to the sentence's local indices.

内容的提问来源于stack exchange,提问作者user17179901

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 20:32:33