CORBA AMI TAO sendc_方法多次调用后冻结问题求助
服务器被SIGSTOP停止后,多次调用sendc_系列方法导致程序冻结
问题场景
- 执行
kill -SIGSTOP命令停止目标服务器 - 多次调用
sendc_EventItem等sendc_系列异步CORBA方法 - 预期:
sendc_EventItem方法正常返回控制权 - 实际:程序出现冻结现象
冻结时的堆栈信息
Stack: __kernel_vsyscall 0x00000000f7f38b19 pthread_sigmask 0x00000000f7ee2f42 [Inlined] ACE_OS::thr_sigsetmask(int, const __sigset_t *, __sigset_t *) OS_NS_Thread.inl:3023 ACE_Sig_Guard::~ACE_Sig_Guard() Signal.cpp:49 ACE_Select_Reactor_T::any_ready(ACE_Select_Reactor_Handle_Set &) Select_Reactor_T.cpp:46 ACE_Select_Reactor_T::wait_for_multiple_events(ACE_Select_Reactor_Handle_Set &, ACE_Time_Value *) Select_Reactor_T.cpp:1078 ACE_TP_Reactor::get_event_for_dispatching(ACE_Time_Value *) TP_Reactor.cpp:443 [Inlined] ACE_TP_Reactor::dispatch_i(ACE_Time_Value *, ACE_TP_Token_Guard &) TP_Reactor.cpp:178 ACE_TP_Reactor::handle_events(ACE_Time_Value *) TP_Reactor.cpp:171 [Inlined] ACE_Reactor::handle_events(ACE_Time_Value *) Reactor.inl:179 TAO_ORB_Core::run(ACE_Time_Value *, int) ORB_Core.cpp:2305 CORBA::ORB::perform_work(ACE_Time_Value *) ORB.cpp:208 TAO_Leader_Follower_Flushing_Strategy::flush_transport(TAO_Transport *, ACE_Time_Value *) Leader_Follower_Flushing_Strategy.cpp:57 TAO_Transport::send_asynchronous_message_i(TAO_Stub *, const ACE_Message_Block *, ACE_Time_Value *) Transport.cpp:1618 TAO_Transport::send_message_shared_i(TAO_Stub *, TAO_Message_Semantics, const ACE_Message_Block *, ACE_Time_Value *) Transport.cpp:1383 TAO_Transport::send_message_shared(TAO_Stub *, TAO_Message_Semantics, const ACE_Message_Block *, ACE_Time_Value *) Transport.cpp:274 TAO_IIOP_Transport::send_message(TAO_OutputCDR &, TAO_Stub *, TAO_ServerRequest *, TAO_Message_Semantics, ACE_Time_Value *) IIOP_Transport.cpp:236 TAO_IIOP_Transport::send_request(TAO_Stub *, TAO_ORB_Core *, TAO_OutputCDR &, TAO_Message_Semantics, ACE_Time_Value *) IIOP_Transport.cpp:208 TAO::Remote_Invocation::send_message(TAO_OutputCDR &, TAO_Message_Semantics, ACE_Time_Value *) Remote_Invocation.cpp:165 TAO::Asynch_Remote_Invocation::remote_invocation(ACE_Time_Value *) Asynch_Invocation.cpp:111 TAO::Asynch_Invocation_Adapter::invoke_twoway(TAO_Operation_Details &, TAO_Pseudo_Var_T<…> &, TAO::Profile_Transport_Resolver &, ACE_Time_Value *&, TAO::Invocation_Retry_State *) Asynch_Invocation_Adapter.cpp:192 TAO::Invocation_Adapter::invoke_remote_i(TAO_Stub *, TAO_Operation_Details &, TAO_Pseudo_Var_T<…> &, ACE_Time_Value *&, TAO::Invocation_Retry_State *) Invocation_Adapter.cpp:275 TAO::Invocation_Adapter::invoke_i(TAO_Stub *, TAO_Operation_Details &) Invocation_Adapter.cpp:98 TAO::Invocation_Adapter::invoke(const TAO::Exception_Data *, unsigned long) Invocation_Adapter.cpp:47 TAO::Asynch_Invocation_Adapter::invoke(Messaging::ReplyHandler *, void (*const &)(TAO_InputCDR &, Messaging::ReplyHandler *, unsigned int)) Asynch_Invocation_Adapter.cpp:102 Subscriber::sendc_EventItem(AMI_SubscriberHandler *, const Toris2_Item &, const Toris2_Attrs &) subscriberC.cpp:162 toris2::server::subscriber::Subscribers::Event(const toris2::server::Toris2Item &&, const std::vector<…> &&) subscribers.cpp:74
问题分析
从堆栈调用链可以看出,冻结发生在ACE_Select_Reactor_T::any_ready方法中——这是ACE的Select反应器在等待IO事件的逻辑。
当服务器被SIGSTOP停止后,TCP连接并未断开,但服务器无法处理请求或返回响应。此时调用异步sendc_方法时,TAO的Leader_Follower_Flushing_Strategy会触发ORB的事件循环(CORBA::ORB::perform_work)来尝试刷新传输层数据。由于服务器无响应,Select反应器进入无限等待状态,导致调用线程被阻塞,程序表现为冻结。
解决方案建议
- 设置异步调用超时:在调用
sendc_EventItem时传入合理的ACE_Time_Value超时参数,避免线程无限阻塞 - 调整ORB反应器配置:修改TAO的Reactor超时设置,确保
handle_events不会无限等待;例如为Select Reactor设置默认超时时间 - 分离事件循环线程:单独启动ORB工作线程处理事件循环,避免业务线程因触发事件处理而被阻塞
- 连接状态预检测:在调用sendc_方法前,通过TAO的API检测连接可用性,避免向已无响应的服务器发送请求
内容的提问来源于stack exchange,提问作者Александр Гордиенко
相关产品推荐
相关产品推荐

