Hystrix Feign针对500错误的重试机制失效,求排查帮助
Let's break down why your Feign retry logic isn't kicking in when the TEST-CONFIG service returns a 500 error, and walk through actionable fixes:
1. Hystrix Timeout is Cutting Off Retries Early
Feign's retry process needs to complete before Hystrix triggers a timeout and fallback. By default, Hystrix has a 1-second timeout, which is likely too short for your 5 retry attempts (even with short initial delays). If Hystrix times out first, it will abort the Feign call entirely, skipping retries.
Fix:
Increase the Hystrix timeout to give Feign enough time to complete all retries. Add this to your application.yml or application.properties:
hystrix: command: # Apply to all Hystrix commands, or target your specific client default: execution: isolation: thread: timeoutInMilliseconds: 30000 # 30 seconds to cover multiple retries # Optional: Target only your TestClient's test() method TestClient#test(): execution: isolation: thread: timeoutInMilliseconds: 30000
2. Feign's Retryer Doesn't Work with Async Return Types (Observable)
Your TestClient returns an Observable<String>, which is an asynchronous type. Feign's built-in Retryer is designed for synchronous calls only—it won't automatically retry async streams.
Fixes:
- Option 1: Switch to a synchronous return type (if your use case allows) to leverage Feign's native retry:
@FeignClient(name = "TEST-CONFIG", configuration = FeignRetryConfig.class, fallbackFactory = XYZClientFallbackFactory.class) public interface TestClient { @RequestMapping(value = "/test", method = RequestMethod.GET, consumes = MediaType.APPLICATION_JSON_VALUE) String test(); // Replace Observable<String> with String } - Option 2: Add retry logic directly to the Observable stream using RxJava's
retryWhenoperator:// In your service class where you call TestClient Observable<String> testResponse = testClient.test() .retryWhen(RetryHandler.create(5, 1000)); // Retry 5 times with 1-second delays
3. Hystrix Retries Are Disabled by Default
Hystrix might be blocking retries even if Feign is configured to retry. You need to explicitly enable Hystrix's retry functionality.
Fix:
Add this to your configuration:
hystrix: command: default: execution: retry: enabled: true # Enable Hystrix-level retries fallback: enabled: true # Ensure fallback is enabled (your error shows fallback failed too, so verify fallback logic as well)
4. Your RetryableException May Be Missing Critical Context
Your FeignErrorDecoder returns a RetryableException, but the constructor you're using doesn't pass the request method or retry timestamp. Some Feign versions require this context to recognize the exception as retryable.
Fix:
Update your ErrorDecoder to use a more complete RetryableException constructor:
public class FeignErrorDecoder implements ErrorDecoder { private final ErrorDecoder defaultErrorDecoder = new Default(); private static final Logger log = LoggerFactory.getLogger(FeignErrorDecoder.class); @Override public Exception decode(String methodKey, Response response) { Exception exception = defaultErrorDecoder.decode(methodKey, response); if (response.status() == 500) { log.error(String.format("##### Got %s response from %s #######", response.status(), methodKey)); return new RetryableException( response.status(), exception.getMessage(), response.request().httpMethod(), exception, Date.from(Instant.now().plusSeconds(1)), // Retry after 1 second response.request() ); } return exception; } }
5. Verify Feign Configuration Priority
Double-check that your FeignRetryConfig is actually being used:
- The
@Primaryannotation ensures it takes precedence over other Feign configurations, but confirm no other@Configurationclasses are overriding it. - Ensure your
@FeignClientcorrectly referencesFeignRetryConfig.classin theconfigurationattribute (your code looks correct here, but it's worth a quick check).
Start with adjusting the Hystrix timeout first—it's the most common culprit here. If that doesn't work, test with a synchronous return type to rule out async-related issues.
内容的提问来源于stack exchange,提问作者user280960

