You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python re.sub将文本转小写时如何排除指定字符串?

Hey there! Let's fix that regex issue you're having. The problem with your current script is that it's not properly prioritizing the strings you want to keep intact—instead, it's likely replacing uppercase letters first and then messing up the reserved terms, leading to that weird duplicated output.

Here's a straightforward approach that works by first matching the exact strings you want to preserve, then converting all other uppercase letters to lowercase:

Solution in Python

First, define your list of reserved strings, then build a regex pattern that captures these terms OR any uppercase letter. We'll use a callback function to handle replacements correctly:

import re

# Your list of strings to keep as-is
reserved_terms = ["TEA", "CHI", "I", "@Begin", "@Language", "ENG", "@Participants", "@Media", "@Transcriber", "@Activities", "SBR", "@Comment"]

# Build regex pattern: capture reserved terms first, then any uppercase letter
# We escape each term to handle special characters (like @) safely
pattern = re.compile(
    r'(' + '|'.join(re.escape(term) for term in reserved_terms) + r')|([A-Z])'
)

def handle_replacement(match):
    # If we matched a reserved term, return it unchanged
    if match.group(1):
        return match.group(1)
    # Otherwise, convert the uppercase letter to lowercase
    return match.group(2).lower()

# Test with your example input
input_text = "@Activities: SBR"
output_text = pattern.sub(handle_replacement, input_text)
print(output_text)  # Output: @Activities: SBR (matches your expected result!)

How this works

The regex prioritizes matching your reserved terms first—so when it sees @Activities or SBR, it captures those as a group and returns them exactly as they are. Any other uppercase letters (not part of a reserved term) get caught by the second group and converted to lowercase.

If you're using JavaScript

The same logic applies here—we'll use a regex with a replacement callback:

const reservedTerms = ["TEA", "CHI", "I", "@Begin", "@Language", "ENG", "@Participants", "@Media", "@Transcriber", "@Activities", "SBR", "@Comment"];

// Escape special characters in reserved terms for regex safety
const escapedTerms = reservedTerms.map(term => term.replace(/[.*+?^${}()|[\]\\]/g, '\\$&'));
const pattern = new RegExp(`(${escapedTerms.join('|')})|([A-Z])`, 'g');

const inputText = "@Activities: SBR";
const outputText = inputText.replace(pattern, (match, reservedGroup, upperGroup) => {
    return reservedGroup ? reservedGroup : upperGroup.toLowerCase();
});

console.log(outputText); // Output: @Activities: SBR

This approach ensures your reserved strings stay untouched, while all other uppercase letters are converted to lowercase without duplication or errors.

内容的提问来源于stack exchange,提问作者Rowland

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:31:49