关于Elasticsearch标准分词器数字加圆点时分词行为的疑问
Hey there! Let’s unpack why the standard tokenizer is behaving this way with your examples—it all boils down to the Unicode Word Boundary rules that drive its logic, specifically how it handles dots between different character types like letters and numbers.
1. system.exe → Single token: system.exe
The dot here is categorized as a MidNumLet character (a punctuation mark designed to connect letters or numbers). Since it’s sandwiched between two letter sequences (system and exe), the tokenizer follows the rule that MidNumLet characters don’t create boundaries between letters. So the entire string stays intact as one token.
2. system32.exe → Tokens: system32 and exe (I suspect your observation of system instead of system32 might be a typo—let’s focus on the core split logic here)
The critical rule at play is Unicode’s WB15 specification. This rule states that while MidNumLet characters (like dots) usually connect letters or connect numbers, they do create a boundary if they come right after a number and right before a letter.
In system32.exe, the segment before the dot ends with a number (32), and the segment after starts with letters (exe). This triggers WB15: the dot is no longer treated as a connector, so the tokenizer splits the string at the dot.
As a quick side note: letters and numbers don’t split by default (per rule WB8), so system32 should always stay as a single token. If you’re seeing system split out, double-check your test setup—there might be an extra step or misconfiguration at play!
3. system32tm.exe → Single token: system32tm.exe
Here, the sequence before the dot is system32tm—a mix of letters, numbers, and more letters. Since letters and numbers don’t create boundaries (WB8), this entire chunk is treated as a single "hybrid" sequence. The dot now sits between a letter (m from tm) and letters (exe), so it goes back to acting as a connector. No split happens, hence the whole string is one token.
The key takeaway is that the dot’s behavior isn’t just about the dot itself—it’s about the context of what’s immediately before and after it. When it’s between letters (or a letter/number sequence that ends with a letter), it binds the parts together; when it’s between a pure number ending and letters, it becomes a split point.
内容的提问来源于stack exchange,提问作者ernirulez

