You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

寻求满足特定分词需求的Token过滤器正则表达式方案

分词器需求与解决方案

需求说明

  • 为包含普通文本、代码及非语言片段的文档构建索引,分词需拆分为小写自然语言字符串和单字符符号
  • 示例:输入 "a few words. Cost*count",需生成令牌 [a] [few] [words] [.] [cost] [*] [count]
  • 复合词处理规则:-、_、. 位于两个单词字符之间时,作为复合词连接符;否则作为单独符号

现有实现

已通过PatternTokenizer实现基础分词逻辑,代码如下:

public static final String tokenRgx = "(([A-Za-z0-9]+[-_.])*[A-Za-z0-9]+)|[^A-Za-z0-9\\s]{1}";

protected TokenStreamComponents createComponents(String fieldName) {
  PatternTokenizer src = new PatternTokenizer(Pattern.compile(tokenRgx), 0);
  TokenStream result = new LowerCaseFilter(src);
  return new TokenStreamComponents(src, result);
}

待优化问题

现有分词器可区分句末句号与复合词中的句号,但无法将复合词拆分为保留末尾连接符的片段。例如输入 "a few words. class.simple_method_name. dd-mm-yyyy.",期望生成令牌:
[a] [few] [words] [.] [class.] [simple_] [method_] [name] [.] [dd-] [mm-] [yyyy] [.]
尝试使用PatternCaptureGroupTokenFilter但未找到合适的正则表达式。

最终解决方案

感谢@rici提供的支持小数的正则表达式,可满足上述拆分需求:

String tokenRegex = "-?[0-9]+\\.[0-9]+|[A-Za-z0-9]+([-_.](?=[A-Za-z0-9]))?|[^A-Za-z0-9\\s]";

内容的提问来源于stack exchange,提问作者TrevorN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 13:50:26