如何用Python提取PDF论文的标题、作者及邮箱?
提取PDF论文标题、作者及邮箱的最优方案
我目前通过PyPDF2读取PDF并提取原始文本,已经实现了移除换行、制表符的文本清理函数,也能轻松提取邮箱,但不确定如何高效提取优先级最高的标题和作者信息。以下是现有代码及文本示例,以及针对标题、作者的最优提取方法:
现有基础代码
PDF文本提取代码
from PyPDF2 import PdfReader import re reader = PdfReader("paper.pdf") text = "" for page in reader.pages: text += page.extract_text() + "\n"
文本清理函数
def remove_newlines_tabs(text): """移除文本中的换行、制表符及转义字符""" # 替换各类换行、制表符和转义斜杠,修复可能拆分的邮箱后缀 formatted_text = text.replace('\\n', ' ').replace('\n', ' ').replace('\t',' ').replace('\\', ' ').replace('. com', '.com') return formatted_text
文本示例
- 原始文本:
'Title\nGoes\nHere\nAuthor Name (sdsd@mail.net)\nUniversity of Teeyab\nSeptember 6, 2022\nSome text in the Document.\n'
- 清理后文本:
'Title Goes Here Author Name (sdsd@mail.net) University of Teeyab September 6, 2022 Some text in the Document. '
标题提取方法(优先级最高)
论文标题通常位于PDF最开头的位置,且具有「原始文本中被换行拆分、清理后为连续短语、多为首字母大写或全大写」的特征,以下是两种可靠提取方案:
方案1:结合邮箱定位精准提取
利用邮箱通常在作者信息后的特性,反向定位标题范围:
def extract_title(formatted_text): # 先匹配邮箱位置 email_match = re.search(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b', formatted_text) if email_match: # 找邮箱所在括号的左边界,括号前为作者名,作者名前即为标题 author_end = formatted_text.rfind('(', 0, email_match.start()) if author_end != -1: return formatted_text[:author_end].strip() # 无括号时,取邮箱前的前8个连续词作为标题候选 before_email_words = [word for word in formatted_text[:email_match.start()].split() if word] return ' '.join(before_email_words[:8]) # 未匹配到邮箱时,直接取文本开头前8个词 return ' '.join(formatted_text.split()[:8]) if formatted_text.split() else None
方案2:正则匹配标题格式特征
针对标题首字母大写的共性,用正则匹配开头的连续大写首字母单词:
def extract_title_regex(formatted_text): title_match = re.match(r'^([A-Z][a-zA-Z0-9\-:]+(\s+[A-Z][a-zA-Z0-9\-:]+){2,7})', formatted_text) return title_match.group(1).strip() if title_match else None
作者及邮箱提取方法
邮箱可直接用正则提取,作者名通常位于标题之后、邮箱之前,可结合邮箱位置反向定位:
def extract_author_and_email(formatted_text): emails = re.findall(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b', formatted_text) authors = [] title = extract_title(formatted_text) for email in emails: email_start = formatted_text.find(email) # 定位邮箱所在括号的左边界 left_paren = formatted_text.rfind('(', 0, email_start) if left_paren != -1: author = formatted_text[:left_paren].strip() # 排除标题部分干扰 if title and author.startswith(title): author = author[len(title):].strip() authors.append(author) else: # 无括号时,取邮箱前3个词作为作者名候选 before_email_words = formatted_text[:email_start].split() author = ' '.join(before_email_words[-3:]) if len(before_email_words)>=3 else ' '.join(before_email_words) authors.append(author.strip()) return list(zip(authors, emails)) if authors and emails else []
整合使用示例
# 读取并清理文本 raw_text = text # 此处为从PDF提取的原始文本 clean_text = remove_newlines_tabs(raw_text) # 提取信息 title = extract_title(clean_text) author_email_list = extract_author_and_email(clean_text) # 输出结果 print(f"标题:{title}") print("作者及邮箱:") for author, email in author_email_list: print(f"- {author}: {email}")
注意事项
- 不同期刊PDF格式差异较大,上述方法基于通用论文格式适配,遇到特殊排版需调整正则或规则
- 若PyPDF2文本提取效果不佳,可替换为
pdfplumber,它能更好保留排版结构,提升定位精准度
内容的提问来源于stack exchange,提问作者user309575
相关产品推荐
相关产品推荐

