You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写支持波斯语与英语的JavaScript正则以提取单词首字符

How to Create a Regex for Persian/Mixed-Language Acronyms

Great question! Your current regex works perfectly for English, but it falls short for Persian (and other non-Latin languages) because JavaScript's default regex tokens like \w and \b only recognize ASCII characters—they don't account for Unicode scripts like Persian's Arabic-based letters. Let's build a solution that supports both Persian and mixed-language text.

The Problem with Your Current Regex

  • \w matches only ASCII letters, numbers, and underscores ([A-Za-z0-9_]), so it ignores Persian characters like گ or ج.
  • \b (word boundary) relies on ASCII rules, which don't work reliably for right-to-left (RTL) languages like Persian, leading to incorrect boundary detection.

The Solution: Unicode-Aware Regex

We'll use Unicode property classes and lookbehind assertions to correctly match the first letter of every word, regardless of language.

Step-by-Step Code Example

// Persian-only example
var persianSentence = 'گروه جوانان خلاق';
var persianMatches = persianSentence.match(/(?<!\p{Letter})\p{Letter}/gu);
var persianAcronym = persianMatches.join('');
console.log(persianAcronym); // Output: گجخ

// Mixed-language example (English + Persian)
var mixedSentence = 'Hello World گروه جوانان خلاق';
var mixedMatches = mixedSentence.match(/(?<!\p{Letter})\p{Letter}/gu);
var mixedAcronym = mixedMatches.join('');
console.log(mixedAcronym); // Output: HWگجخ

Breakdown of the Regex

Let's break down what each part does:

  • (?<!\p{Letter}): A negative lookbehind assertion that checks the character before the current position. It ensures we're only matching a letter that isn't preceded by another letter (i.e., the start of a word, or the beginning of the string).
  • \p{Letter}: Matches any Unicode letter—this includes English, Persian, and almost every other language's alphabet. This only works if we add the u (Unicode) flag.
  • g: Global flag to find all matches in the string, not just the first one.
  • u: Unicode flag to enable support for Unicode property classes and correctly handle non-ASCII characters.

Edge Cases to Consider

  • If your text includes numbers or symbols, this regex will ignore them (since \p{Letter} only matches letters). If you want to include numbers in acronyms, adjust the regex to (?<!\p{Letter}|\p{Number})[\p{Letter}\p{Number}].
  • For languages with diacritics (like some Persian variations), \p{Letter} still works, as it includes accented characters by default.

内容的提问来源于stack exchange,提问作者jones

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 02:32:51