You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MySQL无效字符串值验证:如何校验文本框输入的非法字符?

Hey there! Let's break down how to validate user input to catch invalid text or illegal characters—especially when users paste or type in other languages. I'll cover regex solutions, alternative methods, and key considerations based on common use cases.

1. Regex for Blocking Non-Target Languages

If your goal is to restrict input to a specific set of languages (e.g., only Chinese + English + basic punctuation), use a regex that explicitly allows those characters and rejects everything else.

Example 1: Allow Chinese, English, Numbers, and Common Punctuation

^[\u4e00-\u9fa5a-zA-Z0-9\s.,!?'"()<>{}[\]\\/-]*$

Let's break this down:

  • \u4e00-\u9fa5: Matches all simplified Chinese characters
  • a-zA-Z0-9: Matches English letters and digits
  • \s: Allows whitespace (spaces, tabs, newlines—adjust if you want to block newlines)
  • .,!?'"()<>{}[\]\\/-: A list of common punctuation; add/remove based on your needs
  • ^ and $: Ensure the entire input is validated (no hidden invalid characters at the start/end)

Example 2: Allow Only English (No Other Languages)

Simplify the regex if you only want English text:

^[a-zA-Z0-9\s.,!?'"()<>{}[\]\\/-]*$
2. Regex for Blocking Control Characters & Garbage Text

Sometimes users paste invisible control characters (like non-printable ASCII codes) or invalid UTF-8 garbage. This regex allows all "normal" visible text (any language's letters, numbers, punctuation, and spaces) while blocking weird stuff:

^[\p{L}\p{N}\p{P}\p{Z}]*$
  • \p{L}: Matches all Unicode letters (covers every language—Japanese, Arabic, Spanish, etc.)
  • \p{N}: Matches all Unicode digits
  • \p{P}: Matches all Unicode punctuation
  • \p{Z}: Matches all Unicode whitespace

Note: Not all regex engines support Unicode property classes (\p{...}). For JavaScript, you'll need to use a library like XRegExp if you want full Unicode support.

Block Specific Languages (e.g., Ban Japanese/Arabic)

If you need to allow most languages but block specific ones, use a negative match:

^[^\p{Script:Hiragana}\p{Script:Katakana}\p{Script:Arabic}]*$

This blocks Japanese hiragana/katakana and Arabic script. Replace the script names with the ones you want to exclude (e.g., \p{Script:Latin} to block English/other Latin-based languages).

3. Non-Regex Validation Methods

If regex feels too rigid, you can validate character by character using your programming language's built-in tools:

JavaScript Example

Check each character against allowed Unicode ranges or explicit values:

function isValidInput(input) {
  const allowedChars = new Set([
    ' ', '.', ',', '!', '?', '\'', '"', '(', ')', '[', ']', '{', '}', '\\', '/', '-'
  ]);

  for (const char of input) {
    const charCode = char.charCodeAt(0);
    // Allow English letters, digits, Chinese, or allowed punctuation
    const isAllowed = 
      (charCode >= 0x41 && charCode <= 0x5A) || // A-Z
      (charCode >= 0x61 && charCode <= 0x7A) || // a-z
      (charCode >= 0x30 && charCode <= 0x39) || // 0-9
      (charCode >= 0x4E00 && charCode <= 0x9FA5) || // Chinese
      allowedChars.has(char);

    if (!isAllowed) return false;
  }
  return true;
}

Python Example

Use the unicodedata module to check character categories:

import unicodedata

def is_valid_input(input_str):
    for char in input_str:
        # Allow letters (L), numbers (N), punctuation (P), whitespace (Z)
        category = unicodedata.category(char)
        if not category.startswith(('L', 'N', 'P', 'Z')):
            return False
    return True
4. Key Things to Remember
  • Define your business rules first: Are you blocking other languages entirely, or just garbage characters? Don't over-restrict—if your product supports multilingual users, allow all valid languages and only block control characters/garbage.
  • Test edge cases: Don't forget about full-width characters (e.g., 0 instead of 0), emojis, or accented letters (like é or ñ). Adjust your regex/validation logic to include these if needed.
  • Handle copy-paste gracefully: Users often paste text from other apps, which might include hidden formatting characters. Your validation should catch these without being overly strict.

内容的提问来源于stack exchange,提问作者Hisham shahid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:40:44