WSO2 API Manager 3.2.0高可用端点请求仍失败问题
WSO2 APIM负载均衡端点挂起后仍有请求路由失败问题
配置详情
我们已按WSO2 APIM 3.2.0官方文档配置负载均衡类型的API端点高可用,具体配置如下:
EndpointType: Load Balanced Endpoint Suspension State: ErrorCode = Connection Failed and Connection Closed Initial Duration = 1800000 Max Duration = 1800000 Factor = 1 EndPoint Timeout State: Retries Before Suspension = 2 Retry Delay = 15000
问题现象
采用上述配置后,部分请求仍被路由至已处于挂起状态的端点,导致请求失败。
失败响应示例
{ "fault": { "code": 303000, "type": "Status report", "message": "Runtime Error", "description": "Failover endpoint : NewMCMInboundChannel-RESTAPIService--vv2_APIproductionEndpoint - no ready child endpoints" } } { "fault": { "code": 101503, "type": "Status report", "message": "Runtime Error", "description": "Error connecting to the back end" } }
错误日志详情
TID: [-1] [] [2023-01-09 12:25:48,776] WARN {org.apache.synapse.transport.passthru.ConnectCallback} - Connection refused or failed for : /100.66.213.31:7010 TID: [-1234] [] [2023-01-09 12:25:48,777] WARN {org.apache.synapse.endpoints.EndpointContext} - Endpoint : NewMCMInboundChannel-RESTAPIService--vv2_APIproductionEndpoint_0 with address http://100.66.213.31:7010/mcm-provider is marked as TIMEOUT and will be retried : 1 more time/s after : Mon Jan 09 12:26:03 UTC 2023 until its marked SUSPENDED for failure TID: [-1234] [] [2023-01-09 12:25:48,777] WARN {org.apache.synapse.endpoints.LoadbalanceEndpoint} - Endpoint [NewMCMInboundChannel-RESTAPIService--vv2_APIproductionEndpoint] Detect a Failure in a child endpoint : Endpoint [NewMCMInboundChannel-RESTAPIService--vv2_APIproductionEndpoint_0] TID: [-1234] [] [2023-01-09 12:25:48,778] INFO {org.apache.synapse.mediators.builtin.LogMediator} - {api:admin--NewMCMInboundChannel-RESTAPIService:vv2} STATUS = Executing default 'fault' sequence, ERROR_CODE = 101503, ERROR_MESSAGE = Error connecting to the back end TID: [-1] [] [2023-01-09 12:27:26,156] WARN {org.apache.synapse.transport.passthru.ConnectCallback} - Connection refused or failed for : /100.66.213.31:7010 TID: [-1234] [] [2023-01-09 12:27:26,158] WARN {org.apache.synapse.endpoints.EndpointContext} - Endpoint : NewMCMInboundChannel-RESTAPIService--vv2_APIproductionEndpoint_0 with address http://100.66.213.31:7010/mcm-provider is marked as TIMEOUT and will be retried : 0 more time/s after : Mon Jan 09 12:27:41 UTC 2023 until its marked SUSPENDED for failure TID: [-1234] [] [2023-01-09 12:27:26,158] WARN {org.apache.synapse.endpoints.LoadbalanceEndpoint} - Endpoint [NewMCMInboundChannel-RESTAPIService--vv2_APIproductionEndpoint] Detect a Failure in a child endpoint : Endpoint [NewMCMInboundChannel-RESTAPIService--vv2_APIproductionEndpoint_0] TID: [-1234] [] [2023-01-09 12:27:26,158] INFO {org.apache.synapse.mediators.builtin.LogMediator} - {api:admin--NewMCMInboundChannel-RESTAPIService:vv2} STATUS = Executing default 'fault' sequence, ERROR_CODE = 101503, ERROR_MESSAGE = Error connecting to the back end TID: [-1] [] [2023-01-09 12:31:36,455] WARN {org.apache.synapse.transport.passthru.ConnectCallback} - Connection refused or failed for : /100.66.213.31:7010 TID: [-1234] [] [2023-01-09 12:31:36,459] INFO {org.apache.synapse.endpoints.EndpointContext} - Endpoint : NewMCMInboundChannel-RESTAPIService--vv2_APIproductionEndpoint_0 with address http://100.66.213.31:7010/mcm-provider has been marked for SUSPENSION, but no further retries remain. Thus it will be SUSPENDED. TID: [-1234] [] [2023-01-09 12:31:36,459] WARN {org.apache.synapse.endpoints.EndpointContext} - Suspending endpoint : NewMCMInboundChannel-RESTAPIService--vv2_APIproductionEndpoint_0 with address http://100.66.213.31:7010/mcm-provider - current suspend duration is : 1800000ms - Next retry after : Mon Jan 09 13:01:36 UTC 2023 TID: [-1234] [] [2023-01-09 12:31:36,460] WARN {org.apache.synapse.endpoints.LoadbalanceEndpoint} - Endpoint [NewMCMInboundChannel-RESTAPIService--vv2_APIproductionEndpoint] Detect a Failure in a child endpoint : Endpoint [NewMCMInboundChannel-RESTAPIService--vv2_APIproductionEndpoint_0] TID: [-1234] [] [2023-01-09 12:31:36,460] INFO {org.apache.synapse.mediators.builtin.LogMediator} - {api:admin--NewMCMInboundChannel-RESTAPIService:vv2} STATUS = Executing default 'fault' sequence, ERROR_CODE = 101503, ERROR_MESSAGE = Error connecting to the back end
问题分析与解决建议
- 缩短端点挂起前的重试周期:从日志可见,端点从标记超时到最终挂起存在较长等待窗口,这段时间内请求仍会被路由到故障端点。可减少
Retries Before Suspension次数,或缩短Retry Delay,让端点更快进入挂起状态。 - 完善错误码匹配范围:当前配置仅针对
Connection Failed and Connection Closed触发挂起,但日志中出现的是连接拒绝(Connection refused),需将该场景对应的错误码(如101503)加入匹配列表,确保所有后端连接失败场景都能触发端点挂起。 - 确认负载均衡路由策略:检查负载均衡算法配置,确保算法会跳过未就绪(TIMEOUT状态)的端点,避免请求被路由到半故障状态的节点。
- 配置无可用端点兜底策略:针对
303000错误(无就绪子端点),可添加备用端点或在API故障序列中自定义处理逻辑,避免返回原生错误信息。
内容的提问来源于stack exchange,提问作者PradeepKumar
相关产品推荐
相关产品推荐

