Lucene StandardQueryParser错误解析日期字符串的正确配置方法
根因分析
你为DATE类型PointField配置PointsConfig时使用了仅支持数字解析的DecimalFormat,它遇到ISO8601格式的日期字符串(如2021-06-04T18:00:00Z)时,会在第一个非数字字符-处截断,仅提取前四位数字2021作为解析结果,这就是范围查询被错误处理为date:[2021 TO 2021]的原因。
解决方案
自定义一个支持ISO8601日期格式解析的Format实现,替换DATE类型对应的PointsConfig里的DecimalFormat即可。
步骤1:实现自定义日期解析类
import java.text.Format; import java.text.ParsePosition; import java.util.Date; import org.apache.solr.common.util.DateUtil; public class ISO8601ToEpochFormat extends Format { // 时间戳精度配置:秒级时间戳填1000,毫秒级填1 private static final long TIMESTAMP_SCALE = 1000; @Override public Object parseObject(String source, ParsePosition pos) { try { // 复用Solr内置的日期解析能力,兼容多种ISO8601变种格式 Date date = DateUtil.parseDate(source.trim()); pos.setIndex(source.length()); return date.getTime() / TIMESTAMP_SCALE; } catch (Exception e) { pos.setErrorIndex(0); return null; } } @Override public StringBuffer format(Object obj, StringBuffer toAppendTo, java.text.FieldPosition pos) { if (obj instanceof Long timestamp) { return toAppendTo.append(new Date(timestamp * TIMESTAMP_SCALE).toInstant().toString()); } return toAppendTo.append(obj); } }
步骤2:修改DATE类型的PointsConfig配置
修改extractPointsConfig方法中DATE分支的逻辑:
private PointsConfig extractPointsConfig(FieldType type) { switch (type.getNumberType()) { case DATE: // 日期类型使用自定义的ISO8601转时间戳解析器 return new PointsConfig(new ISO8601ToEpochFormat(), Long.class); case LONG: return new PointsConfig(new DecimalFormat(), Long.class); case INTEGER: return new PointsConfig(new DecimalFormat(), Integer.class); case FLOAT: return new PointsConfig(new DecimalFormat(), Float.class); case DOUBLE: return new PointsConfig(new DecimalFormat(), Double.class); } return null; }
注意事项
- 请根据你Solr中存储的日期时间戳精度调整
TIMESTAMP_SCALE的取值,保证解析出的时间戳和存储值一致。 - 你当前配置的
KeywordTokenizerFactory是正确的,能保证完整的日期字符串被传入解析器,不会被分词截断。
内容的提问来源于stack exchange,提问作者wero026
相关产品推荐
相关产品推荐

