Python/Pandas下含文本数字日期的文件名标准化处理咨询
Hey there! Let's work through your filename standardization problem together. I notice you're new to Python and Pandas, and you've hit two specific snags with your current code—let's break them down and get your filenames looking exactly like your examples.
Your Key Challenges
- Adding an underscore before date strings (like
20200418) - Preserving abbreviations like
ARAasArainstead of splitting them intoA_R_A
What's Wrong with the Original Code?
Your current change_case function iterates character-by-character and adds an underscore every time it hits an uppercase letter. That's why ARA gets turned into A_R_A—it sees each uppercase 'A' as a new "word" to split. Also, it doesn't handle date formatting at all.
The Solution: Token-Based Processing
Instead of splitting by individual uppercase letters, we'll split the filename into words/tokens (using spaces as separators), process each token separately, then join them back with underscores. This fixes both your issues in one go.
Step-by-Step Modified Code
First, we'll write helper functions to handle different types of tokens (abbreviations, dates, numbered labels, regular words), then apply this to all files in your folder:
import os def process_token(token): # Keep 8-digit dates exactly as they are if token.isdigit() and len(token) == 8: return token # Handle labeled versions like V34 or V3 (keep number, capitalize first letter) if any(char.isdigit() for char in token): return token[0].upper() + token[1:] # Convert all-uppercase abbreviations (like ARA) to title case (Ara) if token.isupper() and len(token) > 1: return token.capitalize() # Convert regular words to title case (e.g., "inoc" → "Inoc") return token.capitalize() def standardize_filename(filename): # Split filename from its extension (e.g., ".xlsx") name_part, extension = os.path.splitext(filename) # Split the name into tokens, removing any extra spaces tokens = [token.strip() for token in name_part.split() if token.strip()] # Process each token to match your desired format processed_tokens = [process_token(token) for token in tokens] # Join tokens with underscores and reattach the extension return '_'.join(processed_tokens) + extension # Process all files in your target folder folder_path = "C:\\Users\\t\\Documents\\DummyData\\" for filename in os.listdir(folder_path): file_path = os.path.join(folder_path, filename) # Skip folders, only process files if os.path.isfile(file_path): new_filename = standardize_filename(filename) # Only rename if the filename actually changed (avoids errors) if new_filename != filename: os.rename(file_path, os.path.join(folder_path, new_filename)) print(f"Original: {filename} → Standardized: {new_filename}")
How This Fixes Your Issues
- Date Underscores: Since we split the filename by spaces, dates (which are separate tokens) will automatically be preceded by an underscore when we join everything back together.
- Preserving Abbreviations: By splitting into tokens first, we treat
ARAas a single unit. Theprocess_tokenfunction converts it to title case (Ara) instead of splitting it into individual letters.
Testing with Your Examples
Let's check your sample filenames:
- Original:
ARA Inoc Start Times V34 20200418 .xlsx→ Standardized:Ara_Inoc_Start_Times_V34_20200418.xlsx - Original:
Batch Start Time V3 20200418.xlsx→ Standardized:Batch_Start_Time_V3_20200418.xlsx
Perfect match for your desired output!
Extra Notes
- If your dates ever appear attached to other text (e.g.,
Batch20200418), we could add a regex step to extract the date, but your examples show dates as separate tokens, so this code works as-is. - The code skips renaming files that already match the standardized format to avoid unnecessary operations.
内容的提问来源于stack exchange,提问作者derek9988

