如何使用Python实现学术论文PDF文件的命名格式批量转换
Python自动化重命名学术论文PDF文件
Hey there! I totally get that wrapping your head around Python automation for this specific PDF renaming task can feel overwhelming when you're new to coding. Let's walk through this step by step with a straightforward, easy-to-follow solution.
First, let's recap the requirements clearly
- Original filename format:
Last1, F., Last2, F., & Last3, F. (YYYY). Paper Title Here. Journal Name
Example:Cresswell, K., Worth, A., & Sheikh, A. (2011). Implementing and adopting electronic health record systems. Clinical governance- an international journal - Target filename format:
FirstAuthorLast_FirstAuthorInitial_YYYY_FirstThreeTitleWords
Example:Cresswell_K_2011_Implementing and adopting
The Python solution
We'll use Python's built-in os module to handle file system operations, and regular expressions (re) to extract the exact parts we need from the filenames. Here's the full code with detailed comments to help you follow along:
import os import re def rename_pdf(filename): # Skip any non-PDF files in the folder if not filename.endswith(".pdf"): return filename # Step 1: Extract first author's last name and initial # Matches patterns like "Cresswell, K.," at the start of the filename author_match = re.search(r"^([A-Za-z]+),\s*([A-Za-z])\.,", filename) if not author_match: # If we can't parse the author section, return the original name to avoid errors return filename last_name = author_match.group(1) initial = author_match.group(2) # Step 2: Extract the 4-digit publication year (pattern like "(2011)") year_match = re.search(r"\((\d{4})\)", filename) year = year_match.group(1) if year_match else "UnknownYear" # Step 3: Extract the first three words of the paper title # The title sits between the year's closing parenthesis/period and the journal name's leading period title_part = re.search(r"\(\d{4}\)\.\s*(.*?)\.\s*[A-Z]", filename) if not title_part: title_words = ["UnknownTitle"] else: title_text = title_part.group(1) # Split the title into words, take the first 3, then join back with spaces title_words = title_text.split()[:3] title_str = " ".join(title_words) # Combine all parts into the new filename, keep the .pdf extension new_filename = f"{last_name}_{initial}_{year}_{title_str}.pdf" # Remove any invalid characters (like slashes, colons) that would break the filename new_filename = re.sub(r'[<>:"/\\|?*]', '', new_filename) return new_filename # Apply this function to all PDFs in your target folder # Replace this path with the actual folder path containing your PDFs folder_path = "/path/to/your/pdf/folder" for filename in os.listdir(folder_path): old_path = os.path.join(folder_path, filename) # Skip subdirectories, only process files if not os.path.isfile(old_path): continue new_filename = rename_pdf(filename) new_path = os.path.join(folder_path, new_filename) # Only rename if the new name is different from the original if old_path != new_path: os.rename(old_path, new_path) print(f"Renamed: {filename} -> {new_filename}")
Let's break down what each part does
- Author extraction: The regex grabs the first author's last name (text before the first comma) and their first initial (the letter after the comma, before the period).
- Year extraction: We look for a 4-digit number wrapped in parentheses—this is the standard publication year format in academic filenames.
- Title extraction: We isolate the text between the year section and the journal name, split it into individual words, and keep just the first three.
- Filename cleanup: We strip out any characters that aren't allowed in file names (like
<,>,:) to prevent system errors. - Bulk processing: We loop through every file in your target folder, apply the renaming logic, and update the file names only if a valid new name is generated.
Quick tips for using this code
- Replace
/path/to/your/pdf/folderwith your actual PDF directory path (on Windows, this might look likeC:\\Users\\YourName\\Documents\\AcademicPapers). - Always test on a copy of your PDFs first! It's smart to make a backup before bulk-renaming files to avoid accidental data loss.
- If you run into edge cases (like authors with hyphenated last names), you can tweak the regex patterns to match your specific filename variations.
内容的提问来源于stack exchange,提问作者Jus
相关产品推荐
相关产品推荐

