如何编写支持波斯语与英语的JavaScript正则以提取单词首字符
How to Create a Regex for Persian/Mixed-Language Acronyms
Great question! Your current regex works perfectly for English, but it falls short for Persian (and other non-Latin languages) because JavaScript's default regex tokens like \w and \b only recognize ASCII characters—they don't account for Unicode scripts like Persian's Arabic-based letters. Let's build a solution that supports both Persian and mixed-language text.
The Problem with Your Current Regex
\wmatches only ASCII letters, numbers, and underscores ([A-Za-z0-9_]), so it ignores Persian characters likeگorج.\b(word boundary) relies on ASCII rules, which don't work reliably for right-to-left (RTL) languages like Persian, leading to incorrect boundary detection.
The Solution: Unicode-Aware Regex
We'll use Unicode property classes and lookbehind assertions to correctly match the first letter of every word, regardless of language.
Step-by-Step Code Example
// Persian-only example var persianSentence = 'گروه جوانان خلاق'; var persianMatches = persianSentence.match(/(?<!\p{Letter})\p{Letter}/gu); var persianAcronym = persianMatches.join(''); console.log(persianAcronym); // Output: گجخ // Mixed-language example (English + Persian) var mixedSentence = 'Hello World گروه جوانان خلاق'; var mixedMatches = mixedSentence.match(/(?<!\p{Letter})\p{Letter}/gu); var mixedAcronym = mixedMatches.join(''); console.log(mixedAcronym); // Output: HWگجخ
Breakdown of the Regex
Let's break down what each part does:
(?<!\p{Letter}): A negative lookbehind assertion that checks the character before the current position. It ensures we're only matching a letter that isn't preceded by another letter (i.e., the start of a word, or the beginning of the string).\p{Letter}: Matches any Unicode letter—this includes English, Persian, and almost every other language's alphabet. This only works if we add theu(Unicode) flag.g: Global flag to find all matches in the string, not just the first one.u: Unicode flag to enable support for Unicode property classes and correctly handle non-ASCII characters.
Edge Cases to Consider
- If your text includes numbers or symbols, this regex will ignore them (since
\p{Letter}only matches letters). If you want to include numbers in acronyms, adjust the regex to(?<!\p{Letter}|\p{Number})[\p{Letter}\p{Number}]. - For languages with diacritics (like some Persian variations),
\p{Letter}still works, as it includes accented characters by default.
内容的提问来源于stack exchange,提问作者jones
相关产品推荐
相关产品推荐

