Spring State Machine中DefaultStateMachineExecutor延迟事件列表丢失问题
I’ve worked with Spring State Machine in distributed Spring Cloud Stream setups before, so let’s dig into this delayed event list loss issue you’re facing. First, let’s recap your scenario to make sure I’m aligned:
Our typical state machine workflow monitors Spring Batch processing executed by Spring Cloud Stream microservices. We've set up an
ItemStreamwith source, process, and sink microservices that handle batch reading, processing, and writing of records respectively. Between these stream operations, we trigger REST calls containing single events that are consumed by the state machine—but we're encountering an issue where the delayed event list maintained byDefaultStateMachineExecutoris getting lost.
Why This Happens (Common Causes)
- In-Memory Store Limitations: The default
DefaultStateMachineExecutoruses an in-memory delayed event store. In your distributed Spring Cloud Stream setup, if a microservice instance restarts, scales down, or fails, all in-memory delayed events are gone. This is the most likely culprit for your scenario. - Unconfigured Persistence for State Machine: Without a persistent state store (like JDBC, Redis, or MongoDB), the entire state machine state—including delayed events—doesn’t survive restarts or cross-instance handoffs.
- Task Scheduler Shutdown Behavior: If your executor’s task scheduler isn’t configured to wait for delayed tasks on shutdown, events can be dropped when the service restarts or redeploys.
- Concurrency Race Conditions: When submitting events via REST calls concurrently, unsynchronized event handling could lead to events being accidentally removed from the delayed list.
Fixes Tailored to Your Setup
1. Switch to a Persistent Delayed Event Store
This is the critical fix for distributed environments. Replace the in-memory store with a persistent option. For example, using Redis (common in Spring Cloud setups):
First, add the necessary dependency (if you don’t have it already):
<dependency> <groupId>org.springframework.statemachine</groupId> <artifactId>spring-statemachine-redis</artifactId> </dependency>
Then configure the persistent delayed event store and state machine persister:
@Configuration public class StateMachineConfig extends StateMachineConfigurerAdapter<States, Events> { @Autowired private RedisConnectionFactory redisConnectionFactory; @Bean public DelayedEventStore delayedEventStore() { return new RedisDelayedEventStore(redisConnectionFactory); } @Bean public StateMachinePersister<States, Events> stateMachinePersister() { return new RedisStateMachinePersister<>(redisConnectionFactory); } @Override public void configure(StateMachineConfigurationConfigurer<States, Events> config) throws Exception { config .withConfiguration() .persister(stateMachinePersister()); } }
2. Configure a Resilient Task Scheduler
Make sure your task scheduler waits for in-flight delayed events on shutdown to avoid dropping them:
@Bean public TaskScheduler stateMachineTaskScheduler() { ThreadPoolTaskScheduler scheduler = new ThreadPoolTaskScheduler(); scheduler.setPoolSize(4); // Adjust based on your workload scheduler.setThreadNamePrefix("sm-executor-"); scheduler.setWaitForTasksToCompleteOnShutdown(true); scheduler.setAwaitTerminationSeconds(60); // Give time for delayed events to finish or persist return scheduler; }
Then link this scheduler to your state machine executor:
@Override public void configure(StateMachineConfigurationConfigurer<States, Events> config) throws Exception { config .withConfiguration() .persister(stateMachinePersister()) .taskScheduler(stateMachineTaskScheduler()); }
3. Validate Concurrency Controls
If you’re submitting events concurrently via REST, ensure your event submission logic is thread-safe. Use the state machine’s built-in sendEvent method, which is designed to handle concurrent calls, but avoid custom wrapping that could introduce race conditions.
4. Verify Event Delay Logic
Double-check that the delay you’re setting on events doesn’t exceed the maximum retention time of your persistent store. For example, if using Redis, ensure key expiration isn’t configured to delete delayed event keys before they’re supposed to be processed.
Final Thoughts
In your distributed Spring Cloud Stream + Spring Batch setup, in-memory state management is a non-starter for durable event handling. Persistence is the foundation here—without it, you’ll keep losing delayed events whenever instances change state. Take the time to test the persistent setup with restarts and scale events to confirm the issue is resolved.
内容的提问来源于stack exchange,提问作者Seek Ecom

