You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python实现学术论文PDF文件的命名格式批量转换

Python自动化重命名学术论文PDF文件

Hey there! I totally get that wrapping your head around Python automation for this specific PDF renaming task can feel overwhelming when you're new to coding. Let's walk through this step by step with a straightforward, easy-to-follow solution.

First, let's recap the requirements clearly

  • Original filename format: Last1, F., Last2, F., & Last3, F. (YYYY). Paper Title Here. Journal Name
    Example: Cresswell, K., Worth, A., & Sheikh, A. (2011). Implementing and adopting electronic health record systems. Clinical governance- an international journal
  • Target filename format: FirstAuthorLast_FirstAuthorInitial_YYYY_FirstThreeTitleWords
    Example: Cresswell_K_2011_Implementing and adopting

The Python solution

We'll use Python's built-in os module to handle file system operations, and regular expressions (re) to extract the exact parts we need from the filenames. Here's the full code with detailed comments to help you follow along:

import os
import re

def rename_pdf(filename):
    # Skip any non-PDF files in the folder
    if not filename.endswith(".pdf"):
        return filename
    
    # Step 1: Extract first author's last name and initial
    # Matches patterns like "Cresswell, K.," at the start of the filename
    author_match = re.search(r"^([A-Za-z]+),\s*([A-Za-z])\.,", filename)
    if not author_match:
        # If we can't parse the author section, return the original name to avoid errors
        return filename
    last_name = author_match.group(1)
    initial = author_match.group(2)
    
    # Step 2: Extract the 4-digit publication year (pattern like "(2011)")
    year_match = re.search(r"\((\d{4})\)", filename)
    year = year_match.group(1) if year_match else "UnknownYear"
    
    # Step 3: Extract the first three words of the paper title
    # The title sits between the year's closing parenthesis/period and the journal name's leading period
    title_part = re.search(r"\(\d{4}\)\.\s*(.*?)\.\s*[A-Z]", filename)
    if not title_part:
        title_words = ["UnknownTitle"]
    else:
        title_text = title_part.group(1)
        # Split the title into words, take the first 3, then join back with spaces
        title_words = title_text.split()[:3]
    title_str = " ".join(title_words)
    
    # Combine all parts into the new filename, keep the .pdf extension
    new_filename = f"{last_name}_{initial}_{year}_{title_str}.pdf"
    
    # Remove any invalid characters (like slashes, colons) that would break the filename
    new_filename = re.sub(r'[<>:"/\\|?*]', '', new_filename)
    
    return new_filename

# Apply this function to all PDFs in your target folder
# Replace this path with the actual folder path containing your PDFs
folder_path = "/path/to/your/pdf/folder"

for filename in os.listdir(folder_path):
    old_path = os.path.join(folder_path, filename)
    # Skip subdirectories, only process files
    if not os.path.isfile(old_path):
        continue
    
    new_filename = rename_pdf(filename)
    new_path = os.path.join(folder_path, new_filename)
    
    # Only rename if the new name is different from the original
    if old_path != new_path:
        os.rename(old_path, new_path)
        print(f"Renamed: {filename} -> {new_filename}")

Let's break down what each part does

  1. Author extraction: The regex grabs the first author's last name (text before the first comma) and their first initial (the letter after the comma, before the period).
  2. Year extraction: We look for a 4-digit number wrapped in parentheses—this is the standard publication year format in academic filenames.
  3. Title extraction: We isolate the text between the year section and the journal name, split it into individual words, and keep just the first three.
  4. Filename cleanup: We strip out any characters that aren't allowed in file names (like <, >, :) to prevent system errors.
  5. Bulk processing: We loop through every file in your target folder, apply the renaming logic, and update the file names only if a valid new name is generated.

Quick tips for using this code

  • Replace /path/to/your/pdf/folder with your actual PDF directory path (on Windows, this might look like C:\\Users\\YourName\\Documents\\AcademicPapers).
  • Always test on a copy of your PDFs first! It's smart to make a backup before bulk-renaming files to avoid accidental data loss.
  • If you run into edge cases (like authors with hyphenated last names), you can tweak the regex patterns to match your specific filename variations.

内容的提问来源于stack exchange,提问作者Jus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 20:12:42