You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取PDF论文的标题、作者及邮箱?

提取PDF论文标题、作者及邮箱的最优方案

我目前通过PyPDF2读取PDF并提取原始文本,已经实现了移除换行、制表符的文本清理函数,也能轻松提取邮箱,但不确定如何高效提取优先级最高的标题和作者信息。以下是现有代码及文本示例,以及针对标题、作者的最优提取方法:

现有基础代码

PDF文本提取代码

from PyPDF2 import PdfReader
import re

reader = PdfReader("paper.pdf")
text = ""
for page in reader.pages:
    text += page.extract_text() + "\n"

文本清理函数

def remove_newlines_tabs(text):
    """移除文本中的换行、制表符及转义字符"""
    # 替换各类换行、制表符和转义斜杠,修复可能拆分的邮箱后缀
    formatted_text = text.replace('\\n', ' ').replace('\n', ' ').replace('\t',' ').replace('\\', ' ').replace('. com', '.com')
    return formatted_text

文本示例

  • 原始文本:
'Title\nGoes\nHere\nAuthor Name (sdsd@mail.net)\nUniversity of Teeyab\nSeptember 6, 2022\nSome text in the Document.\n'
  • 清理后文本:
'Title Goes Here Author Name (sdsd@mail.net) University of Teeyab September 6, 2022 Some text in the Document. '

标题提取方法(优先级最高)

论文标题通常位于PDF最开头的位置,且具有「原始文本中被换行拆分、清理后为连续短语、多为首字母大写或全大写」的特征,以下是两种可靠提取方案:

方案1:结合邮箱定位精准提取

利用邮箱通常在作者信息后的特性,反向定位标题范围:

def extract_title(formatted_text):
    # 先匹配邮箱位置
    email_match = re.search(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b', formatted_text)
    if email_match:
        # 找邮箱所在括号的左边界,括号前为作者名,作者名前即为标题
        author_end = formatted_text.rfind('(', 0, email_match.start())
        if author_end != -1:
            return formatted_text[:author_end].strip()
        # 无括号时,取邮箱前的前8个连续词作为标题候选
        before_email_words = [word for word in formatted_text[:email_match.start()].split() if word]
        return ' '.join(before_email_words[:8])
    # 未匹配到邮箱时,直接取文本开头前8个词
    return ' '.join(formatted_text.split()[:8]) if formatted_text.split() else None

方案2:正则匹配标题格式特征

针对标题首字母大写的共性,用正则匹配开头的连续大写首字母单词:

def extract_title_regex(formatted_text):
    title_match = re.match(r'^([A-Z][a-zA-Z0-9\-:]+(\s+[A-Z][a-zA-Z0-9\-:]+){2,7})', formatted_text)
    return title_match.group(1).strip() if title_match else None

作者及邮箱提取方法

邮箱可直接用正则提取,作者名通常位于标题之后、邮箱之前,可结合邮箱位置反向定位:

def extract_author_and_email(formatted_text):
    emails = re.findall(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b', formatted_text)
    authors = []
    title = extract_title(formatted_text)
    
    for email in emails:
        email_start = formatted_text.find(email)
        # 定位邮箱所在括号的左边界
        left_paren = formatted_text.rfind('(', 0, email_start)
        if left_paren != -1:
            author = formatted_text[:left_paren].strip()
            # 排除标题部分干扰
            if title and author.startswith(title):
                author = author[len(title):].strip()
            authors.append(author)
        else:
            # 无括号时,取邮箱前3个词作为作者名候选
            before_email_words = formatted_text[:email_start].split()
            author = ' '.join(before_email_words[-3:]) if len(before_email_words)>=3 else ' '.join(before_email_words)
            authors.append(author.strip())
    
    return list(zip(authors, emails)) if authors and emails else []

整合使用示例

# 读取并清理文本
raw_text = text  # 此处为从PDF提取的原始文本
clean_text = remove_newlines_tabs(raw_text)

# 提取信息
title = extract_title(clean_text)
author_email_list = extract_author_and_email(clean_text)

# 输出结果
print(f"标题:{title}")
print("作者及邮箱:")
for author, email in author_email_list:
    print(f"- {author}: {email}")

注意事项

  • 不同期刊PDF格式差异较大,上述方法基于通用论文格式适配,遇到特殊排版需调整正则或规则
  • 若PyPDF2文本提取效果不佳,可替换为pdfplumber,它能更好保留排版结构,提升定位精准度

内容的提问来源于stack exchange,提问作者user309575

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 12:03:30