如何基于邮箱本地名过滤剩余邮箱地址中的垃圾邮箱?
Great question—filtering spammy local parts is tricky because there’s no one-size-fits-all rule, but there are several practical, layered approaches you can combine to weed out addresses like YSXCVZXFASDFKLS@gmail.com or pp3@gmail.com while keeping valid ones like ebcom@gmail.com.
Start with targeted regex rules to catch obvious spam patterns in local parts:
- Long all-uppercase random strings: Normal users rarely use 8+ character all-caps jumbles. Use this regex to flag them:
^[A-Z]{8,}$ - Short repetitive/sequential patterns: Catch addresses like
pp3orxxx2with rules like^([a-z0-9])\1{1,}\d?$(matches 2+ repeated characters followed by an optional digit) - Vowel-free long strings: Most valid local parts include vowels. Flag strings longer than 5 characters with no vowels (adjust this if you need to accommodate names like
lynch):^[^aeiouAEIOU]{6,}$
Example implementation in Python:
import re def is_spam_local_part(local_part): # Flag long all-uppercase jumbles if re.match(r'^[A-Z]{8,}$', local_part): return True # Flag short repetitive patterns (e.g., pp3, xxx2) if re.match(r'^([a-z0-9])\1{1,}\d?$', local_part): return True # Flag long vowel-free strings if not re.search(r'[aeiouAEIOU]', local_part) and len(local_part) > 5: return True return False
Random spam strings have higher entropy (more "chaos") than meaningful local parts. Calculate entropy to separate random jumbles from valid addresses:
import math from collections import Counter def calculate_entropy(s): if not s: return 0 char_counts = Counter(s) probabilities = [count / len(s) for count in char_counts.values()] return -sum(p * math.log2(p) for p in probabilities) # Example threshold: Entropy > 4.5 + length > 6 = likely spam def is_high_entropy_spam(local_part): entropy = calculate_entropy(local_part) return entropy > 4.5 and len(local_part) > 6
For context: YSXCVZXFASDFKLS will have an entropy near 4.7, while ebcom sits around 2.3.
Spam local parts are almost always unique (one-off random strings), while valid addresses often repeat across your dataset (e.g., info@company.com, john.doe@domain.net):
- Filter out local parts that appear only once (balance this with a whitelist of common valid names like
admin,support) - Prioritize retaining local parts that appear 2+ times across different domains
With 400k+ addresses, a lightweight ML model will scale better than manual rules:
- Extract features: Length, entropy, vowel count, repeated character ratio, presence of digits, and whether the string matches common dictionary words
- Label a small subset: Manually tag 1-2k addresses as spam/valid
- Train a classifier: Use models like Naive Bayes, Random Forest, or XGBoost—these work well with tabular features and don’t require massive compute
- Don’t block short local parts entirely (e.g.,
jo@gmail.comis valid) - Whitelist common business-focused local names (
contact,sales,billing) - Combine rules instead of using them in isolation (e.g., only flag a string if it has high entropy and is all caps and has no vowels)
内容的提问来源于stack exchange,提问作者Matheus Leonardo

