Shell脚本正则表达式无法限制带重音字符问题排查
Ah, I see the issue here! Your regex ^[a-zA-Z0-9]+$ isn't blocking accented strings like Léo because of bash's locale-dependent regex behavior. Let me break this down and show you how to fix it.
Why Your Current Regex Fails
When your system uses a UTF-8 locale (like en_US.UTF-8, which is common these days), bash expands the [a-zA-Z] character class to include all Unicode letter characters—not just the ASCII A-Z/a-z range. That means accented characters like é get treated as valid letters, so Léo slips past your check.
Solution 1: Force ASCII Mode with Locale
The simplest fix is to temporarily switch bash to the C locale (pure ASCII mode) when defining and using your regex. This ensures [a-zA-Z0-9] only matches ASCII letters and numbers, blocking any accented or non-ASCII characters.
Here's how to update your script:
# Switch to ASCII locale to enforce strict character matching LC_ALL=C regex="^[a-zA-Z0-9]+$" while read -r name; do if [[ "$name" =~ $regex ]]; then echo "Valid name: $name" else echo "Invalid name (contains accented/non-ASCII characters): $name" fi done
Solution 2: Use Explicit ASCII Range or POSIX Class
You can also use an explicit ASCII range or POSIX character class (again, paired with LC_ALL=C for consistency):
LC_ALL=C # Option 1: Explicit ASCII letter/number range regex="^[0-9A-Za-z]+$" # Option 2: POSIX alnum class (only matches ASCII chars in C locale) regex="^[[:alnum:]]+$"
Test It Out
Let's verify this works as expected:
# Test with accented string echo "Léo" | ./your_script.sh # Output: Invalid name (contains accented/non-ASCII characters): Léo # Test with valid ASCII string echo "Leo123" | ./your_script.sh # Output: Valid name: Leo123
内容的提问来源于stack exchange,提问作者danaso

