基于Spacy NER实体标签过滤训练数据:移除无实体句子的实现问题
Got it, let's tackle this problem step by step. The core challenge here is twofold: splitting your text into sentences based on periods and newlines, then filtering out sentences that don't contain any entities—and adjusting the entity character indices since they're tied to the original full text, not the split sentences.
Here's a complete implementation that builds on your existing code:
from tqdm import tqdm import spacy from spacy.tokens import DocBin def split_text_into_sentences_with_indices(text): """Split text into sentences using . and \n as delimiters, return each sentence with its original start/end indices.""" sentences = [] current_start = 0 i = 0 while i < len(text): # Check for sentence delimiters if text[i] in ('.', '\n'): # Extract the sentence (trim whitespace to avoid empty sentences) sentence_text = text[current_start:i].strip() if sentence_text: sentences.append((sentence_text, current_start, i)) # Move past the delimiter current_start = i + 1 i += 1 # Add the last sentence if there's any remaining text final_sentence = text[current_start:].strip() if final_sentence: sentences.append((final_sentence, current_start, len(text))) return sentences nlp = spacy.blank("en") db = DocBin() train_data = [('Christmas Perot 2021 TSO\nSkip to Main Content HOME CONCERTS EVENTS ABOUT STAFF EDUCATION SUPPORT US More Use tab to navigate through the menu items. BUY TICKETS SUNDAY, DECEMBER 12, 2021 I PEROT THEATRE I 4:00 PM\nPOPS I Christmas at The Perot\nCLICK HERE to purchase tickets, or contact the Texarkana Symphony Orchestra at 870.773.3401\nA Texarkana Tradition Join the TSO, the Texarkana Jazz Orchestra, and the TSO Chamber Singers, for this holiday concert for the whole family.\nDon’t miss seeing the winner of TSO’s 11th Annual Celebrity Conductor Competition\nBack to Events 2019 Texarkana Symphony Orchestra', {'entities': [(375, 399, 'organization'), (290, 318, 'organization'), (220, 242, 'production_name'), (169, 186, 'performance_date'), (189, 202, 'auditorium'), (205, 212, 'performance_starttime'), (409, 428, 'organization')]})] for text, annot in tqdm(train_data): # Split text into sentences with their original positions sentences_with_indices = split_text_into_sentences_with_indices(text) for sent_text, sent_start, sent_end in sentences_with_indices: # Check if any entity falls within this sentence's bounds sent_entities = [] for ent_start, ent_end, ent_label in annot["entities"]: # Entity is fully contained within the sentence if sent_start <= ent_start < ent_end <= sent_end: # Adjust entity indices to be relative to the sentence start rel_start = ent_start - sent_start rel_end = ent_end - sent_start sent_entities.append((rel_start, rel_end, ent_label)) # Only process sentences that have at least one entity if sent_entities: doc = nlp.make_doc(sent_text) ents = [] for start, end, label in sent_entities: span = doc.char_span(start, end, label=label, alignment_mode="contract") if span is None: print(f"Skipping entity ({start}, {end}, {label}) in sentence: {sent_text}") else: ents.append(span) doc.ents = ents db.add(doc) # Save the processed DocBin if needed # db.to_disk("./filtered_ner_data.spacy")
Let's break down what this does:
Sentence Splitting with Indices:
- The
split_text_into_sentences_with_indicesfunction splits your text using.and\nas delimiters, and keeps track of each sentence's starting and ending character positions in the original full text. This is crucial for matching entities to sentences. - It also trims whitespace to avoid empty sentences from consecutive delimiters.
- The
Filtering & Index Adjustment:
- For each sentence, we check if any of the original entities are fully contained within the sentence's bounds.
- If an entity belongs to the sentence, we subtract the sentence's starting index from the entity's original start/end to get relative indices for the sentence.
- We only create a
Docobject for sentences that have at least one entity, skipping the ones with no entities entirely.
Integration with spaCy:
- The rest of the code follows your original workflow: creating
Docobjects, adding adjusted entities, and populating theDocBin.
- The rest of the code follows your original workflow: creating
Testing with your sample data:
The first line Christmas Perot 2021 TSO has no entities, so it gets skipped entirely. Sentences containing entities like BUY TICKETS SUNDAY, DECEMBER 12, 2021 I PEROT THEATRE I 4:00 PM will be retained, with their entities adjusted to the sentence's local indices.
内容的提问来源于stack exchange,提问作者user17179901

