Spacy处理德语维基分句:速度慢且效果不佳的技术咨询
Hi there! Let’s tackle your two big issues with spaCy for German Wikipedia sentence splitting—glacial processing speed and wonky split results. Here’s how to fix both:
Your current setup is loading the full spaCy pipeline (tagger, parser, NER, etc.) even though you only need tokenization and sentence splitting. Plus, processing files one-by-one without batch optimizations is killing your throughput. Try these tweaks:
Strip Unnecessary Pipeline Components
Load only the parts you need to cut down on overhead:nlp = spacy.load('de_core_news_sm', disable=['tagger', 'parser', 'ner', 'lemmatizer']) # Add the sentencizer explicitly since we disabled the parser (which usually handles it) nlp.add_pipe('sentencizer')Use
nlp.pipe()for Batch Processingnlp.pipe()is optimized for handling multiple texts at once, and reading files in one go is more efficient than line-by-line concatenation:for file in files: file_path = join(rootdir, path, file) new_file_name = join(output_dir, file + ".txt") # Read entire file at once with open(file_path, 'r', encoding='utf-8') as f: content = f.read() # Process with batch optimization for doc in nlp.pipe([content], batch_size=50): # Adjust batch size based on your RAM with open(new_file_name, 'w', encoding='utf-8') as new_file: for sent in doc.sents: new_file.write(sent.text.strip() + '\n') new_file.write('\n')Go Multiprocessing for Massive File Counts
With 55k files, parallel processing will slash your runtime. Use Python’smultiprocessingto split the work across your CPU cores:from multiprocessing import Pool def process_single_file(file): # Load the model inside the worker process (avoids cross-process issues) nlp = spacy.load('de_core_news_sm', disable=['tagger', 'parser', 'ner', 'lemmatizer']) nlp.add_pipe('sentencizer') file_path = join(rootdir, path, file) new_file_name = join(output_dir, file + ".txt") with open(file_path, 'r', encoding='utf-8') as f: content = f.read() doc = nlp(content) with open(new_file_name, 'w', encoding='utf-8') as new_file: for sent in doc.sents: new_file.write(sent.text.strip() + '\n') new_file.write('\n') if __name__ == '__main__': # Use 4 processes (adjust based on your CPU core count) with Pool(processes=4) as pool: pool.map(process_single_file, files)
The small de_core_news_sm model often struggles with German-specific edge cases like name abbreviations or complex date formats. Try these fixes:
Upgrade to a Larger Model
Switch tode_core_news_lg(larger rule-based model) orde_core_news_trf(transformer-based) for better sentence boundary detection:nlp = spacy.load('de_core_news_lg', disable=['tagger', 'parser', 'ner', 'lemmatizer']) nlp.add_pipe('sentencizer')Add Custom Boundary Rules
For cases likeMaria I. (England)where the abbreviation shouldn’t trigger a sentence split, define custom logic to override the default behavior:from spacy.lang.de import German from spacy.pipeline import Sentencizer nlp = German() # Initialize sentencizer with standard punctuation sentencizer = Sentencizer(nlp.vocab, punct_chars=["!", ".", "?", "…"]) nlp.add_pipe(sentencizer) # Custom rule: Don't start a new sentence after "I." if it's followed by a proper noun def adjust_sentence_boundaries(doc): for token in doc[:-1]: if token.text == "I." and doc[token.i + 1].pos_ == "PROPN": doc[token.i + 1].is_sent_start = False return doc nlp.add_pipe(adjust_sentence_boundaries, after='sentencizer')Double-Check Input Data
In your example, the output is missing18.from* 18. Februar 1516—that looks like either a rare spaCy bug or a typo/corruption in the original file. Verify the input content to rule out data issues first.
Give these changes a shot—you should see a massive speed improvement and much cleaner sentence splits.
内容的提问来源于stack exchange,提问作者maggie ezzat

