寻求满足特定分词需求的Token过滤器正则表达式方案
分词器需求与解决方案
需求说明
- 为包含普通文本、代码及非语言片段的文档构建索引,分词需拆分为小写自然语言字符串和单字符符号
- 示例:输入
"a few words. Cost*count",需生成令牌[a] [few] [words] [.] [cost] [*] [count] - 复合词处理规则:
-、_、.位于两个单词字符之间时,作为复合词连接符;否则作为单独符号
现有实现
已通过PatternTokenizer实现基础分词逻辑,代码如下:
public static final String tokenRgx = "(([A-Za-z0-9]+[-_.])*[A-Za-z0-9]+)|[^A-Za-z0-9\\s]{1}"; protected TokenStreamComponents createComponents(String fieldName) { PatternTokenizer src = new PatternTokenizer(Pattern.compile(tokenRgx), 0); TokenStream result = new LowerCaseFilter(src); return new TokenStreamComponents(src, result); }
待优化问题
现有分词器可区分句末句号与复合词中的句号,但无法将复合词拆分为保留末尾连接符的片段。例如输入 "a few words. class.simple_method_name. dd-mm-yyyy.",期望生成令牌:[a] [few] [words] [.] [class.] [simple_] [method_] [name] [.] [dd-] [mm-] [yyyy] [.]
尝试使用PatternCaptureGroupTokenFilter但未找到合适的正则表达式。
最终解决方案
感谢@rici提供的支持小数的正则表达式,可满足上述拆分需求:
String tokenRegex = "-?[0-9]+\\.[0-9]+|[A-Za-z0-9]+([-_.](?=[A-Za-z0-9]))?|[^A-Za-z0-9\\s]";
内容的提问来源于stack exchange,提问作者TrevorN
相关产品推荐
相关产品推荐

