Boost.Asio异步执行器析构时偶发SIGSEGV崩溃排查求助
问题:Boost.Asio异步执行器析构时偶发SIGSEGV崩溃
实现了一个基于Boost.Asio的异步执行器类,支持指定延迟后在另一线程执行回调,还能重置延迟(看门狗功能)。多数情况运行正常,但约10%概率在析构时触发SIGSEGV崩溃。原以为析构前已等待线程join,this指针应始终有效,但错误栈显示this=0x0,无法定位根源,需解决析构时无法正确取消后台任务导致的崩溃问题。
变量类型定义
boost::asio::chrono::milliseconds delay_ms_; boost::asio::io_context io_; boost::asio::steady_timer timer_; boost::thread thread_; std::mutex mutable mutex_; bool is_running = false;
类实现代码
AsyncExecutor::AsyncExecutor(std::chrono::milliseconds const &delay_ms, Callback callback) : delay_ms_(delay_ms), callback_(callback), io_(), timer_(io_, delay_ms_) { } AsyncExecutor::~AsyncExecutor() { // Make sure the timer stops when the object is destroyed & callback is not called. stop(); } void AsyncExecutor::start() { std::lock_guard<std::mutex> lock(mutex_); if (is_running) return; is_running = true; if (io_.stopped()) { io_.reset(); // Reset io_service if it has stopped } timer_.expires_after(delay_ms_); timer_.async_wait(boost::bind(&AsyncExecutor::timer_ran, this, boost::asio::placeholders::error)); // Start the io context in a separate thread. // Once the timer fires, the thread will exit. thread_ = boost::thread(boost::bind(&boost::asio::io_service::run, &io_)); } void AsyncExecutor::stop() { { std::lock_guard<std::mutex> lock(mutex_); if (!is_running) return; // Cancel the timer. boost::asio::post(io_, [this]() { this->timer_.cancel(); }); } // Finaly wait for the thread to finish. if (this->thread_.joinable()) this->thread_.join(); std::lock_guard<std::mutex> lock(mutex_); is_running = false; } void AsyncExecutor::reset() { std::lock_guard<std::mutex> lock(mutex_); if (!is_running) return; // Reset the timer. boost::asio::post(io_, [this]() { this->timer_.expires_after(this->delay_ms_); this->timer_.async_wait(boost::bind(&AsyncExecutor::timer_ran, this, boost::asio::placeholders::error)); }); } void AsyncExecutor::timer_ran(const boost::system::error_code &ec) { { std::lock_guard<std::mutex> lock(mutex_); is_running = false; } // If the timer was cancelled, we don't want to execute the callback or detach the thread. // Reset also calls this function with ec, so we want to keep the thread attached. // This complicates things & means that join() must be called in stop() to wait for the thread to finish. // But if the thread fires the callback, it will detach itself. if (ec) { return; } if (this->callback_) this->callback_(); io_.stop(); // Detach the thread, as it is no longer needed. It will be cleaned up by the OS. this->thread_.detach(); } bool AsyncExecutor::is_timer_running() const { std::lock_guard<std::mutex> lock(mutex_); return is_running; }
GDB栈追踪信息
[gdb-7] [New Thread 0x7fff1a7fc000 (LWP 385144)] [gdb-7] [Thread 0x7fff1a7fc000 (LWP 385075) exited] [gdb-7] [gdb-7] Thread 171 "scout_control_u" received signal SIGSEGV, Segmentation fault. [gdb-7] [Switching to Thread 0x7fff2253c000 (LWP 385067)] [gdb-7] boost::asio::detail::epoll_reactor::run (this=0x0, usec=<optimized out>, ops=...) at /usr/include/boost/asio/detail/impl/epoll_reactor.ipp:462 [gdb-7] 462 if (timer_fd_ == -1) [gdb-7] #0 boost::asio::detail::epoll_reactor::run(long, boost::asio::detail::op_queue<boost::asio::detail::scheduler_operation>&) [gdb-7] (this=0x0, usec=<optimized out>, ops=...) [gdb-7] at /usr/include/boost/asio/detail/impl/epoll_reactor.ipp:462 [gdb-7] #1 0x00007ffff6f0aef3 in boost::asio::detail::scheduler::do_run_one(boost::asio::detail::conditionally_enabled_mutex::scoped_lock&, boost::asio::detail::scheduler_thread_info&, boost::system::error_code const&) [gdb-7] (this=this@entry=0x7fff1412e5f0, lock=..., this_thread=..., ec=...) [gdb-7] at /usr/include/boost/asio/detail/impl/scheduler.ipp:465 [gdb-7] #2 0x00007ffff6f0bb01 in boost::asio::detail::scheduler::run(boost::system::error_code&) (this=0x7fff1412e5f0, ec=...) [gdb-7] at /usr/include/boost/asio/detail/impl/scheduler.ipp:204 [gdb-7] #3 0x00007ffff6f0d552 in boost::asio::io_context::run() (this=<optimized out>) [gdb-7] at /usr/include/boost/asio/impl/io_context.ipp:63 [gdb-7] #4 0x00007ffff1f570cb in () [gdb-7] at /lib/x86_64-linux-gnu/libboost_thread.so.1.74.0 [gdb-7] #5 0x00007ffff7094ac3 in start_thread (arg=<optimized out>) [gdb-7] at ./nptl/pthread_create.c:442 [gdb-7] #6 0x00007ffff7126850 in clone3 () [gdb-7] at ../sysdeps/unix/sysv/linux/x86_64/clone3.S:81
问题根源与修复方案
核心问题分析
- 线程detach导致悬空引用:
timer_ran中定时器正常触发时调用thread_.detach(),但析构时stop()仍尝试对已detach的线程执行join(),属于未定义行为。更严重的是,detach后的线程可能仍在运行io_context::run(),此时AsyncExecutor对象已被析构,io_成员已销毁,线程访问已销毁的io_context内部结构(如epoll_reactor)就会触发SIGSEGV,对应栈中this=0x0的情况。 - 线程生命周期管理混乱:
start()中每次启动都会重新赋值thread_,若之前的线程未正确终止,会导致线程对象的未定义行为。同时timer_ran直接detach线程,破坏了stop()通过join等待线程结束的逻辑,导致析构时线程可能仍在运行,访问已销毁的对象成员。
修复步骤
- 移除线程detach操作:取消
timer_ran中的thread_.detach(),统一由stop()负责join线程,确保线程在对象析构前完全终止。 - 修正io_context停止逻辑:在
stop()中直接调用io_.stop()终止事件循环,确保线程能快速退出,无需通过post异步取消定时器。 - 同步线程启动状态:
start()中确保之前的线程已完全终止后再启动新线程,避免线程对象重复赋值问题。 - 统一互斥锁保护:操作
thread_、is_running时全程持有互斥锁,避免多线程竞争。
修复后的关键代码示例
修改后的stop()函数
void AsyncExecutor::stop() { std::lock_guard<std::mutex> lock(mutex_); if (!is_running) return; // 直接停止io_context,终止事件循环 io_.stop(); // 取消定时器 timer_.cancel(); // 等待线程结束 if (thread_.joinable()) { thread_.join(); } is_running = false; }
修改后的timer_ran()函数
void AsyncExecutor::timer_ran(const boost::system::error_code &ec) { std::lock_guard<std::mutex> lock(mutex_); is_running = false; if (ec) { return; } if (callback_) { callback_(); } // 不再detach线程,由stop()统一管理线程生命周期 }
修改后的start()函数
void AsyncExecutor::start() { std::lock_guard<std::mutex> lock(mutex_); if (is_running) return; // 确保之前的线程已终止 if (thread_.joinable()) { thread_.join(); } io_.reset(); timer_.expires_after(delay_ms_); timer_.async_wait(boost::bind(&AsyncExecutor::timer_ran, this, boost::asio::placeholders::error)); thread_ = boost::thread(boost::bind(&boost::asio::io_context::run, &io_)); is_running = true; }
额外注意事项
- 所有访问
is_running、thread_、io_的操作必须在互斥锁保护下,避免多线程竞争。 - 析构函数调用
stop()时,虽已持有对象生命周期控制权,但stop()内部的锁仍需保留,以处理其他线程可能的并发调用。
内容的提问来源于stack exchange,提问作者Krzo
相关产品推荐
相关产品推荐

