Spring Data KeyValue如何检测并获知重复键?
Spring Data KeyValue能否检测重复键?
批处理结果
readCount=116361, filterCount=23687, writeCount=92674 readSkipCount=0, writeSkipCount=0, processSkipCount=0, commitCount=4655
已实现操作
我已定义复合键:
public record LocationKey (String countryCode, String locationCode) implements Serializable { }
在ItemProcessor中配置键:
.withKey(new LocationKey(item.countryCode(), item.locationCode()));
在ItemWriter中执行保存:
toWrite.add(chunk.size()); repository.saveAll(chunk).iterator().forEachRemaining((i) -> written.increment());
执行结果
toWrite : 92674
written : 92674
locations in repository : 92571
文档说明
Returns:the saved entities; will never be null. The returned Iterable will have the same size as the Iterable passed as an argument.
(翻译:返回值:保存后的实体;绝不会为null。返回的Iterable与传入的Iterable大小一致。)
实际行为符合文档描述,但我该如何获知哪些是重复键?是否需要更换其他Map实现来实现需求?
使用版本
- Spring Boot 3.4.1
- Spring Data KeyValue 3.4.1
- Spring Batch 5.2.1
解答
Spring Data KeyValue默认使用SimpleKeyValueRepository,底层依赖Map存储数据。调用saveAll时,重复键会直接覆盖原有数据,但默认不会抛出异常或返回重复提示——这就是写入计数和返回实体数量一致,但仓库最终实体数更少的原因。
要检测并识别重复键,可通过以下方式实现:
1. 写入前主动校验重复
在ItemProcessor或ItemWriter中提前校验当前批次的键:
- 批量查询仓库中当前批次所有键对应的实体,对比找出已存在的键;
- 维护内存缓存(如
ConcurrentHashMap)记录已写入的键,处理每个item时检查重复。
示例(ItemProcessor中):
private final LocationRepository repository; private final Set<LocationKey> processedKeys = ConcurrentHashMap.newKeySet(); @Override public Location process(Location item) throws Exception { LocationKey key = new LocationKey(item.countryCode(), item.locationCode()); if (processedKeys.contains(key) || repository.existsById(key)) { log.warn("检测到重复键:{}", key); return null; // 返回null会被Spring Batch计入filterCount } processedKeys.add(key); return item; }
2. 自定义KeyValueTemplate的Map实现
Spring Data KeyValue允许替换底层存储的Map。默认用ConcurrentHashMap,你可以自定义Map,在put操作时检测重复并记录:
- 用装饰器模式包装原有Map,重写
put/putAll方法,检测到重复键时记录日志或抛出异常; - 通过
KeyValueTemplate构造函数传入自定义Map,注入到Repository中。
示例:
@Bean public KeyValueTemplate keyValueTemplate() { Map<Object, Object> map = new DuplicateDetectingMap(new ConcurrentHashMap<>()); return new KeyValueTemplate(new MapKeyValueAdapter(map)); } static class DuplicateDetectingMap<K, V> implements Map<K, V> { private final Map<K, V> delegate; private final Logger log = LoggerFactory.getLogger(DuplicateDetectingMap.class); public DuplicateDetectingMap(Map<K, V> delegate) { this.delegate = delegate; } @Override public V put(K key, V value) { if (delegate.containsKey(key)) { log.warn("检测到重复键,将覆盖原有数据:{}", key); // 若需抛出异常,可改为 throw new DuplicateKeyException("重复键:" + key); } return delegate.put(key, value); } @Override public void putAll(Map<? extends K, ? extends V> m) { m.forEach(this::put); } // 省略其他Map方法的委托实现... }
3. 结合Spring Batch跳过机制
如果自定义Map中抛出DuplicateKeyException,可在Step配置中添加跳过策略,捕获异常并记录:
@Bean public Step locationStep() { return stepBuilderFactory.get("locationStep") .<Location, Location>chunk(1000) .reader(reader()) .processor(processor()) .writer(writer()) .faultTolerant() .skip(DuplicateKeyException.class) .skipLimit(1000) .listener(new SkipListener<Location, Location>() { @Override public void onSkipInWrite(Location item, Throwable t) { log.warn("写入时跳过重复项:{},原因:{}", item, t.getMessage()); } }) .build(); }
总结:无需更换Spring Data KeyValue,通过前置校验、自定义Map实现或结合Spring Batch跳过机制,即可实现重复键的检测与记录。
内容的提问来源于stack exchange,提问作者wims.tijd

