You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写支持日文与英文姓名的多语言正则表达式(JavaScript)

Supporting Both English and Japanese Names with Regex in JavaScript

Great question! Your original regex ^[a-zA-Z\s]+$ works perfectly for English names, but we can expand it to cover the common characters used in Japanese names without breaking existing functionality. Let’s walk through how to build this step by step.

Key Characters in Japanese Names

Japanese names typically use three core character sets, plus optional formatting symbols:

  • Kanji (漢字): The most common form for family/given names, covered by the Unicode range \u4E00-\u9FA5 (basic CJK unified ideographs, which includes nearly all kanji used in everyday names).
  • Hiragana (ひらがな): Used for phonetic readings or casual given names, covered by \u3040-\u309F.
  • Katakana (カタカナ): Often used for foreign-derived names or formal contexts, covered by \u30A0-\u30FF (includes the long dash ー at \u30FC).
  • Middle dot (・): Some Japanese names use this symbol (\u30FB) to separate family and given names.

The Updated Multi-Language Regex

We’ll combine these Japanese character ranges with your original English set, and add the u modifier (critical for proper Unicode handling in JavaScript):

const multiLangNameRegex = /^[a-zA-Z\s\u4E00-\u9FA5\u3040-\u309F\u30A0-\u30FF\u30FB]+$/u;

Let’s break down each component:

  • a-zA-Z: Matches English uppercase/lowercase letters (your original logic)
  • \s: Matches whitespace (works for English name separators; if you only want regular spaces, replace with )
  • \u4E00-\u9FA5: Matches Japanese kanji characters
  • \u3040-\u309F: Matches all hiragana characters (including浊音/半浊音 variants)
  • \u30A0-\u30FF: Matches all katakana characters (including the long dash ー)
  • \u30FB: Matches the middle dot separator for Japanese names
  • u: Enables Unicode mode, ensuring the regex correctly interprets multi-byte Japanese characters (required to avoid broken matches)

Test It with Real-World Cases

Here’s a quick example to verify the regex works across different name formats:

// Test various valid and invalid name cases
const testNames = [
  "Emma Wilson",         // English name ✅
  "山田 太郎",           // Kanji with space ✅
  "スズキ ヒロシ",       // Katakana with space ✅
  "さとう なおみ",       // Hiragana with space ✅
  "佐藤・花子",          // Kanji with middle dot ✅
  "Mike 鈴木",           // Mixed English + Japanese ✅
  "Tanaka123",           // Contains numbers ❌
  "Smith@Doe",           // Contains special characters ❌
  "カナガワ-ケン"        // Uses invalid dash ❌
];

testNames.forEach(name => {
  console.log(`"${name}": ${multiLangNameRegex.test(name)}`);
});

Edge Case Adjustments

  • If you don’t need to support the middle dot separator, simply remove \u30FB from the character set.
  • The u modifier is supported in all modern browsers and Node.js versions (v6+). Support for older environments is rarely needed today, but if required, you’d need a custom fallback solution.
  • This regex allows any combination of allowed characters—if you need stricter rules (e.g., enforce exactly one separator), you’d need to tweak the pattern further, but this covers most common name formats.

内容的提问来源于stack exchange,提问作者MERLIN THOMAS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:40:52