Python单轮替换多正则表达式问题求助
Hey there! I get it—moving from PHP regex to Python can throw a few curveballs, especially when your go-to patterns stop working. Let's break down why your current code isn't doing what you want, and fix it step by step.
The Main Issue: PHP vs Python Regex Case Insensitivity
In PHP, you add /i at the end of a regex pattern to enable case-insensitive matching. But Python's re module doesn't recognize that suffix—it treats /i as part of the pattern itself, which is why your replacements aren't catching anything (since filenames like (Intro) don't have a trailing /i).
Fix 1: Adjust Your Patterns and Use re.IGNORECASE
First, remove the /i suffix from all your regex strings. Then, pass the flags=re.IGNORECASE parameter to re.sub() to get that case-insensitive behavior you need. Here's your updated code:
import re # Don't forget to import the re module—it's easy to miss! # Updated regex patterns (removed /i suffix) regexes = { r'\s(\(|\[)(.*?)Mix(.*?)(\)|\])': r"", r'\s(\(|\[)(.*?)Version(.*?)(\)|\])': r"", r'\s(\(|\[)(.*?)Remix(.*?)(\)|\])': r"", r'\s(\(|\[)(.*?)Extended(.*?)(\)|\])': r"", r'\s\(remix\)': r"", r'\s\(original\)': r"", r'\s\(intro\)': r"", } def multi_replace(regex_dict, text): for pattern, replacement in regex_dict.items(): # Add the ignore case flag here text = re.sub(pattern, replacement, text, flags=re.IGNORECASE) return text filename = "Testing (Intro)" name = multi_replace(regexes, filename) print(name) # Output: Testing
Fix 2: Use Inline Case Insensitivity (Alternative)
If you prefer to keep the case-insensitive flag directly in your regex patterns (more like PHP's style), you can use Python's inline (?i) flag at the start of each pattern. This way you don't need to pass the flags parameter:
import re regexes = { r'(?i)\s(\(|\[)(.*?)Mix(.*?)(\)|\])': r"", r'(?i)\s(\(|\[)(.*?)Version(.*?)(\)|\])': r"", r'(?i)\s(\(|\[)(.*?)Remix(.*?)(\)|\])': r"", r'(?i)\s(\(|\[)(.*?)Extended(.*?)(\)|\])': r"", r'(?i)\s\(remix\)': r"", r'(?i)\s\(original\)': r"", r'(?i)\s\(intro\)': r"", } def multi_replace(regex_dict, text): for pattern, replacement in regex_dict.items(): text = re.sub(pattern, replacement, text) return text filename = "Testing (Intro)" name = multi_replace(regexes, filename) print(name) # Output: Testing
Bonus: Optimize Your Regex for Efficiency
Since many of your patterns follow a similar structure, you can combine them into a single regex to reduce the number of replacements you need to run. This will make your code faster, especially if you're processing lots of filenames:
import re # Combined pattern that matches all your target strings combined_pattern = r'(?i)\s(\(|\[)(.*?)(Mix|Version|Remix|Extended)(.*?)(\)|\])|\s\(remix\)|\s\(original\)|\s\(intro\)' def clean_filename(text): return re.sub(combined_pattern, "", text) filename = "Testing (Intro)" print(clean_filename(filename)) # Output: Testing # Test another example: "DJ XYZ - Track Name (Extended Mix)" print(clean_filename("DJ XYZ - Track Name (Extended Mix)")) # Output: DJ XYZ - Track Name
This works because the | character acts as an OR in regex—so it will match any of the patterns you defined, all in one pass.
内容的提问来源于stack exchange,提问作者DmVinny

