You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于邮箱本地名过滤剩余邮箱地址中的垃圾邮箱?

Great question—filtering spammy local parts is tricky because there’s no one-size-fits-all rule, but there are several practical, layered approaches you can combine to weed out addresses like YSXCVZXFASDFKLS@gmail.com or pp3@gmail.com while keeping valid ones like ebcom@gmail.com.

1. Pattern-Based Regex Filtering

Start with targeted regex rules to catch obvious spam patterns in local parts:

  • Long all-uppercase random strings: Normal users rarely use 8+ character all-caps jumbles. Use this regex to flag them: ^[A-Z]{8,}$
  • Short repetitive/sequential patterns: Catch addresses like pp3 or xxx2 with rules like ^([a-z0-9])\1{1,}\d?$ (matches 2+ repeated characters followed by an optional digit)
  • Vowel-free long strings: Most valid local parts include vowels. Flag strings longer than 5 characters with no vowels (adjust this if you need to accommodate names like lynch): ^[^aeiouAEIOU]{6,}$

Example implementation in Python:

import re

def is_spam_local_part(local_part):
    # Flag long all-uppercase jumbles
    if re.match(r'^[A-Z]{8,}$', local_part):
        return True
    # Flag short repetitive patterns (e.g., pp3, xxx2)
    if re.match(r'^([a-z0-9])\1{1,}\d?$', local_part):
        return True
    # Flag long vowel-free strings
    if not re.search(r'[aeiouAEIOU]', local_part) and len(local_part) > 5:
        return True
    return False
2. Entropy Score Calculation

Random spam strings have higher entropy (more "chaos") than meaningful local parts. Calculate entropy to separate random jumbles from valid addresses:

import math
from collections import Counter

def calculate_entropy(s):
    if not s:
        return 0
    char_counts = Counter(s)
    probabilities = [count / len(s) for count in char_counts.values()]
    return -sum(p * math.log2(p) for p in probabilities)

# Example threshold: Entropy > 4.5 + length > 6 = likely spam
def is_high_entropy_spam(local_part):
    entropy = calculate_entropy(local_part)
    return entropy > 4.5 and len(local_part) > 6

For context: YSXCVZXFASDFKLS will have an entropy near 4.7, while ebcom sits around 2.3.

3. Frequency Analysis

Spam local parts are almost always unique (one-off random strings), while valid addresses often repeat across your dataset (e.g., info@company.com, john.doe@domain.net):

  • Filter out local parts that appear only once (balance this with a whitelist of common valid names like admin, support)
  • Prioritize retaining local parts that appear 2+ times across different domains
4. Machine Learning (For Large-Scale Accuracy)

With 400k+ addresses, a lightweight ML model will scale better than manual rules:

  1. Extract features: Length, entropy, vowel count, repeated character ratio, presence of digits, and whether the string matches common dictionary words
  2. Label a small subset: Manually tag 1-2k addresses as spam/valid
  3. Train a classifier: Use models like Naive Bayes, Random Forest, or XGBoost—these work well with tabular features and don’t require massive compute
5. Edge Cases to Avoid Over-Filtering
  • Don’t block short local parts entirely (e.g., jo@gmail.com is valid)
  • Whitelist common business-focused local names (contact, sales, billing)
  • Combine rules instead of using them in isolation (e.g., only flag a string if it has high entropy and is all caps and has no vowels)

内容的提问来源于stack exchange,提问作者Matheus Leonardo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:31:14