You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写正则校验任意语言字母开头的多语言字符、空格及指定特殊字符

Fixing Regex to Match Strings with Diacritics/Combining Marks

Alright, let's tackle this regex issue! The problem with your original pattern /^[\p{L} -()']+$/ is that it only matches base letter characters (\p{L}), but characters like the ٌ in Castaٌeda are Unicode combining marks—they're not part of the base letter set, so your regex rejects them outright.

The Solution

Update your regex to include Unicode mark characters (\p{M}) and enable Unicode mode with the u modifier (critical for most modern regex engines to handle Unicode property classes properly):

/^[\p{L}\p{M} '()&-]+$/u

Breakdown of the Pattern

  • \p{L}: Matches any base letter character from any language (Latin, Arabic, Cyrillic, etc.)
  • \p{M}: Matches all Unicode mark characters—this includes non-spacing diacritics (like the ٌ in your example), spacing combining marks, and enclosing marks. It covers every type of accent or modifier that attaches to a base letter.
  • '()&-: Your allowed special characters (note: placing - at the end of the character class avoids needing to escape it)
  • ^ & $: Ensures the entire string adheres to the pattern (no invalid characters at start/end)
  • u modifier: Enables Unicode mode, so the regex engine correctly interprets \p{L} and \p{M} instead of treating them as literal characters.

Testing It Out

This pattern will now match:

  • Castaٌeda (with the Arabic combining damma)
  • José (Spanish tilde)
  • São Paulo (Portuguese tilde)
  • الْعَرَبِيَّة (Arabic with combining marks)
  • Any other string starting with a letter, followed by letters, spaces, and your allowed special characters—including those with diacritics.

Notes for Different Regex Engines

  • JavaScript: Must include the u modifier; without it, \p{} syntax won't work.
  • PHP: Use preg_match with the u flag (e.g., preg_match('/^[\p{L}\p{M} \'()&-]+$/u', $string)).
  • Python: Use the re.UNICODE flag or prepend (?u) to the pattern (e.g., re.match(r'(?u)^[\p{L}\p{M} \'()&-]+$', string)).

If you want to be more restrictive (only match non-spacing diacritics), you can replace \p{M} with \p{Mn} (Unicode Non-Spacing Mark), but \p{M} is more inclusive and safer for most use cases.

内容的提问来源于stack exchange,提问作者jack

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:23:32