PDF分句Token列表转编号句/DataFrame及文内引用句筛选问询
Hey there! Based on your workflow—you've already extracted and cleaned text from PDFs into sentence tokens—let's walk through your two options, plus which one makes sense for different use cases.
Option 1: Convert to a DataFrame with Labels (Highly Recommended for Structured Analysis)
Turning your sentence list into a pandas DataFrame is a great move if you want to keep track of all sentences (not just those with citations) and add metadata/annotations for downstream NLP tasks. This gives you flexibility to filter, sort, and analyze your data later on.
Here's a quick code example to implement this:
import pandas as pd import re # Your sample sentence list sentences = ['this is my new project', 'I am very excited about this (Abbasi, 2015)'] # Create DataFrame df = pd.DataFrame({'sentence': sentences}) # Define a regex pattern to detect in-text citations (adjust based on your citation format) citation_pattern = r'\([A-Za-z]+, \d{4}\)' # Matches (Author, Year) format # Add a label column: 1 if citation exists, 0 otherwise df['has_citation'] = df['sentence'].apply(lambda x: 1 if re.search(citation_pattern, x) else 0) # Optional: Extract the citation text itself if you need it df['citation'] = df['sentence'].apply(lambda x: re.findall(citation_pattern, x) if re.search(citation_pattern, x) else None) print(df)
Why this works:
- You retain full context of all sentences, which is useful if you want to compare cited vs non-cited sentences later.
- The
has_citationlabel makes it trivial to filter just the sentences you need withdf[df['has_citation'] == 1]. - Extracting the citation text lets you link sentences to specific references for deeper analysis (e.g., counting how often each paper is cited in your corpus).
Option 2: Directly Extract Sentences with Citations
If your only immediate goal is to work with cited sentences (and you don't need to keep the non-cited ones), filtering directly from your list is more lightweight.
Here's how to do that:
import re sentences = ['this is my new project', 'I am very excited about this (Abbasi, 2015)'] citation_pattern = r'\([A-Za-z]+, \d{4}\)' # Filter sentences with citations cited_sentences = [sent for sent in sentences if re.search(citation_pattern, sent)] # Optional: Convert to numbered list as you mentioned numbered_cited = [f"{i+1}. {sent}" for i, sent in enumerate(cited_sentences)] for item in numbered_cited: print(item)
When to use this:
- You're focused solely on analyzing cited content (e.g., extracting citation contexts for literature review automation).
- You want to minimize memory usage if working with a huge corpus and don't need the non-cited sentences.
Final Recommendation
Go with the DataFrame approach if you plan to do any kind of comparative analysis (e.g., how cited sentences differ from non-cited ones in terms of tone, topic, etc.) or if you need to track metadata alongside your sentences. If you only need the cited sentences for a single, narrow task, direct extraction is fine—but the DataFrame gives you way more flexibility for future NLP work.
内容的提问来源于stack exchange,提问作者Sri Amudha

