如何判断Spacy提取的文本是否为完整句子并过滤标题类内容?
Great question! Working with PDF-extracted text is always messy—headers, footers, table of contents entries, and random chunks are everywhere. Here’s a practical, spaCy-focused approach to separate complete sentences from the noise:
Part 1: How to Check if a spaCy Processed Text is a Complete Sentence
Complete sentences have a core grammatical structure: a subject (noun/noun phrase) and a finite predicate (a verb that conveys a specific tense/action). Here are two reliable ways to verify this with spaCy:
1. Use Syntactic Dependencies & Part-of-Speech Tags
SpaCy’s dependency parser labels key sentence components, which we can leverage to check for mandatory sentence parts:
- Look for a subject tag like
nsubj(nominal subject) ornsubjpass(passive subject) - Look for a finite verb (not an infinitive, gerund, or past participle) marked with
VERB/AUXand tags likeVBD(past tense),VBZ(3rd person singular present), etc.
Here’s a reusable function:
import spacy nlp = spacy.load("en_core_web_sm") def is_complete_sentence(doc): has_valid_subject = False has_finite_verb = False for token in doc: # Check for a clear sentence subject if token.dep_ in ("nsubj", "nsubjpass"): has_valid_subject = True # Check for a finite (tense-specific) verb if token.pos_ in ("VERB", "AUX") and token.tag_ not in ("VB", "VBG", "VBN"): has_finite_verb = True return has_valid_subject and has_finite_verb # Test it out test_samples = [ "Machine Learning Basics", # Header "The model achieved 95% accuracy on the test set.", # Complete sentence "Appendix A: Additional Metrics" # Header ] for sample in test_samples: doc = nlp(sample) print(f"'{sample}' → Complete sentence? {is_complete_sentence(doc)}")
2.辅助判断:Check End Punctuation
Complete sentences typically end with ., ?, or !. This is a weak check on its own (PDFs often have missing/misplaced punctuation), but it can reinforce the syntactic check above. For example:
def has_sentence_ending_punct(doc): return doc.text.strip()[-1] in (".", "?", "!") if doc.text.strip() else False
Part 2: Filtering Headers, Footers, & Other Non-Sentence Content
Once you can identify complete sentences, use these rules to filter out common PDF noise:
1. Rule-Based Filtering for Headers/TOC Entries
- Short, noun-heavy chunks: Headers are usually short (≤8 tokens) and consist mostly of nouns/proper nouns with no verbs. Combine this with the
is_complete_sentencecheck:def is_likely_header(doc): verb_count = sum(1 for t in doc if t.pos_ in ("VERB", "AUX")) noun_count = sum(1 for t in doc if t.pos_ in ("NOUN", "PROPN")) return verb_count == 0 and noun_count >= 1 and len(doc) <= 8 - Table of Contents patterns: TOC entries often have dotted leaders followed by page numbers (e.g., "Chapter 3 ...... 17"). Use regex to catch these:
import re def is_toc_entry(text): return bool(re.search(r"\.{3,}\s*\d+$", text.strip()))
2. Filter Repeating Footers/Headers
Footers (like page numbers, document titles) repeat across pages. Track text frequency—if a chunk appears 3+ times, it’s almost certainly a footer/header:
from collections import defaultdict text_frequency = defaultdict(int) # First pass: count occurrences of each extracted chunk for chunk in all_pdf_chunks: text_frequency[chunk.strip()] += 1 # Second pass: filter out high-frequency chunks filtered_chunks = [chunk for chunk in all_pdf_chunks if text_frequency[chunk.strip()] < 3]
3. Use Position Metadata (If Available)
If your PDF extraction tool (like PyMuPDF, pdfplumber) provides text coordinates, use them to prioritize filtering:
- Headers are usually in the top 10% of the page
- Footers are usually in the bottom 10% of the page
- Skip chunks in these regions unless they pass the complete sentence check
Final Notes
There’s no one-size-fits-all solution—tweak the thresholds (like token length, frequency count) based on your specific PDF type (academic papers, reports, etc.). For OCR’d PDFs, you may need to add extra checks for misrecognized text.
内容的提问来源于stack exchange,提问作者CrabbyPete

