使用Python re.sub将文本转小写时如何排除指定字符串?
Hey there! Let's fix that regex issue you're having. The problem with your current script is that it's not properly prioritizing the strings you want to keep intact—instead, it's likely replacing uppercase letters first and then messing up the reserved terms, leading to that weird duplicated output.
Here's a straightforward approach that works by first matching the exact strings you want to preserve, then converting all other uppercase letters to lowercase:
Solution in Python
First, define your list of reserved strings, then build a regex pattern that captures these terms OR any uppercase letter. We'll use a callback function to handle replacements correctly:
import re # Your list of strings to keep as-is reserved_terms = ["TEA", "CHI", "I", "@Begin", "@Language", "ENG", "@Participants", "@Media", "@Transcriber", "@Activities", "SBR", "@Comment"] # Build regex pattern: capture reserved terms first, then any uppercase letter # We escape each term to handle special characters (like @) safely pattern = re.compile( r'(' + '|'.join(re.escape(term) for term in reserved_terms) + r')|([A-Z])' ) def handle_replacement(match): # If we matched a reserved term, return it unchanged if match.group(1): return match.group(1) # Otherwise, convert the uppercase letter to lowercase return match.group(2).lower() # Test with your example input input_text = "@Activities: SBR" output_text = pattern.sub(handle_replacement, input_text) print(output_text) # Output: @Activities: SBR (matches your expected result!)
How this works
The regex prioritizes matching your reserved terms first—so when it sees @Activities or SBR, it captures those as a group and returns them exactly as they are. Any other uppercase letters (not part of a reserved term) get caught by the second group and converted to lowercase.
If you're using JavaScript
The same logic applies here—we'll use a regex with a replacement callback:
const reservedTerms = ["TEA", "CHI", "I", "@Begin", "@Language", "ENG", "@Participants", "@Media", "@Transcriber", "@Activities", "SBR", "@Comment"]; // Escape special characters in reserved terms for regex safety const escapedTerms = reservedTerms.map(term => term.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')); const pattern = new RegExp(`(${escapedTerms.join('|')})|([A-Z])`, 'g'); const inputText = "@Activities: SBR"; const outputText = inputText.replace(pattern, (match, reservedGroup, upperGroup) => { return reservedGroup ? reservedGroup : upperGroup.toLowerCase(); }); console.log(outputText); // Output: @Activities: SBR
This approach ensures your reserved strings stay untouched, while all other uppercase letters are converted to lowercase without duplication or errors.
内容的提问来源于stack exchange,提问作者Rowland

