You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何判断Spacy提取的文本是否为完整句子并过滤标题类内容?

How to Identify Complete Sentences & Filter Non-Sentence Content (Headers, Footers, etc.) with spaCy for PDF-Extracted Text

Great question! Working with PDF-extracted text is always messy—headers, footers, table of contents entries, and random chunks are everywhere. Here’s a practical, spaCy-focused approach to separate complete sentences from the noise:


Part 1: How to Check if a spaCy Processed Text is a Complete Sentence

Complete sentences have a core grammatical structure: a subject (noun/noun phrase) and a finite predicate (a verb that conveys a specific tense/action). Here are two reliable ways to verify this with spaCy:

1. Use Syntactic Dependencies & Part-of-Speech Tags

SpaCy’s dependency parser labels key sentence components, which we can leverage to check for mandatory sentence parts:

  • Look for a subject tag like nsubj (nominal subject) or nsubjpass (passive subject)
  • Look for a finite verb (not an infinitive, gerund, or past participle) marked with VERB/AUX and tags like VBD (past tense), VBZ (3rd person singular present), etc.

Here’s a reusable function:

import spacy
nlp = spacy.load("en_core_web_sm")

def is_complete_sentence(doc):
    has_valid_subject = False
    has_finite_verb = False
    
    for token in doc:
        # Check for a clear sentence subject
        if token.dep_ in ("nsubj", "nsubjpass"):
            has_valid_subject = True
        # Check for a finite (tense-specific) verb
        if token.pos_ in ("VERB", "AUX") and token.tag_ not in ("VB", "VBG", "VBN"):
            has_finite_verb = True
    
    return has_valid_subject and has_finite_verb

# Test it out
test_samples = [
    "Machine Learning Basics",  # Header
    "The model achieved 95% accuracy on the test set.",  # Complete sentence
    "Appendix A: Additional Metrics"  # Header
]
for sample in test_samples:
    doc = nlp(sample)
    print(f"'{sample}' → Complete sentence? {is_complete_sentence(doc)}")

2.辅助判断:Check End Punctuation

Complete sentences typically end with ., ?, or !. This is a weak check on its own (PDFs often have missing/misplaced punctuation), but it can reinforce the syntactic check above. For example:

def has_sentence_ending_punct(doc):
    return doc.text.strip()[-1] in (".", "?", "!") if doc.text.strip() else False

Part 2: Filtering Headers, Footers, & Other Non-Sentence Content

Once you can identify complete sentences, use these rules to filter out common PDF noise:

1. Rule-Based Filtering for Headers/TOC Entries

  • Short, noun-heavy chunks: Headers are usually short (≤8 tokens) and consist mostly of nouns/proper nouns with no verbs. Combine this with the is_complete_sentence check:
    def is_likely_header(doc):
        verb_count = sum(1 for t in doc if t.pos_ in ("VERB", "AUX"))
        noun_count = sum(1 for t in doc if t.pos_ in ("NOUN", "PROPN"))
        return verb_count == 0 and noun_count >= 1 and len(doc) <= 8
    
  • Table of Contents patterns: TOC entries often have dotted leaders followed by page numbers (e.g., "Chapter 3 ...... 17"). Use regex to catch these:
    import re
    def is_toc_entry(text):
        return bool(re.search(r"\.{3,}\s*\d+$", text.strip()))
    

2. Filter Repeating Footers/Headers

Footers (like page numbers, document titles) repeat across pages. Track text frequency—if a chunk appears 3+ times, it’s almost certainly a footer/header:

from collections import defaultdict

text_frequency = defaultdict(int)

# First pass: count occurrences of each extracted chunk
for chunk in all_pdf_chunks:
    text_frequency[chunk.strip()] += 1

# Second pass: filter out high-frequency chunks
filtered_chunks = [chunk for chunk in all_pdf_chunks if text_frequency[chunk.strip()] < 3]

3. Use Position Metadata (If Available)

If your PDF extraction tool (like PyMuPDF, pdfplumber) provides text coordinates, use them to prioritize filtering:

  • Headers are usually in the top 10% of the page
  • Footers are usually in the bottom 10% of the page
  • Skip chunks in these regions unless they pass the complete sentence check

Final Notes

There’s no one-size-fits-all solution—tweak the thresholds (like token length, frequency count) based on your specific PDF type (academic papers, reports, etc.). For OCR’d PDFs, you may need to add extra checks for misrecognized text.

内容的提问来源于stack exchange,提问作者CrabbyPete

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:30:20