.NET中AggregateException未被捕获致应用崩溃问题求助
问题背景
应用在VPN断开测试场景下崩溃:连接VPN启动应用后,运行时断开VPN,随即抛出未被捕获的AggregateException导致应用崩溃,且调用栈中无我方代码。核心代码结构如下:
- 仓储方法
GetBatchAsync()在Polly的retryPolicy内被调用 - 重试策略嵌套在
Parallel.ForAsync()中 - 已确认问题与"Enable Just My Code"调试设置无关
相关代码
Parallel.ForAsync逻辑
await Parallel.ForAsync(0, totalBatchesCount, parallelOptions, async (i, cancellationToken) => { //... await retryPolicy.ExecuteAsync(async () => { try { var batch = await _userRepository.GetBatchAsync(startId, endId); //... } catch (Exception ex) { throw new Exception($"Batch failed for startId: {startId} and endId {endId}", ex); } }); });
仓储方法
public async Task<List<User>> GetBatchAsync(int startId, int endId) { try { await using var context = await _contextFactory.CreateDbContextAsync(); return await context.Users .Where(u => u.UserId >= startId && u.UserId < endId) .Include(u => u.Invoice) .AsNoTracking() .ToListAsync(); } catch (AggregateException ex) { Console.WriteLine(ex.ToString()); throw; } catch (Exception ex) { Console.WriteLine(ex.ToString()); throw; } }
Polly重试策略
var retryPolicy = Policy .Handle<Exception>() .WaitAndRetryAsync(POLLY_RETRY_COUNT, retryAttempt => TimeSpan.FromSeconds(Math.Pow(2, retryAttempt)), (exception, timeSpan, retryCount, context) => { _logger.LogWarning($"Retry {retryCount} due to: {exception.Message}, stack trace: {exception.StackTrace}, inner: {exception?.InnerException?.Message}"); });
异常信息
Unhandled exception. System.AggregateException: One or more errors occurred. (Timeout expired. The timeout period elapsed prior to completion of the operation or the server is not responding) ---> MySql.Data.MySqlClient.MySqlException (0x80004005): Timeout expired. The timeout period elapsed prior to completion of the operation or the server is not responding ---> System.TimeoutException: The operation has timed out. at MySql.Data.Common.StreamCreator.<>c.<GetTcpStreamAsync>b__8_1() at System.Threading.CancellationTokenSource.Invoke(Delegate d, Object state, CancellationTokenSource source) at System.Threading.ExecutionContext.RunInternal(ExecutionContext executionContext, ContextCallback callback, Object state) --- End of stack trace from previous location --- at System.Threading.ExecutionContext.RunInternal(ExecutionContext executionContext, ContextCallback callback, Object state) at System.Threading.CancellationTokenSource.ExecuteCallbackHandlers(Boolean throwOnFirstException) --- End of inner exception stack trace --- at System.Threading.CancellationTokenSource.ExecuteCallbackHandlers(Boolean throwOnFirstException) at System.Threading.TimerQueueTimer.Fire(Boolean isThreadPool) at System.Threading.TimerQueue.FireNextTimers() at System.Threading.ThreadPoolWorkQueue.Dispatch() at System.Threading.PortableThreadPool.WorkerThread.WorkerThreadStart()
问题分析
这个未捕获异常的根源在于MySQL客户端底层的CancellationToken回调抛出了异常,该异常发生在ThreadPool线程上,脱离了业务代码的async/await上下文,因此现有try-catch块、Polly重试策略都无法捕获它。
从异常栈可以看到,异常起源于MySql.Data.Common.StreamCreator.GetTcpStreamAsync的超时逻辑,触发CancellationTokenSource.ExecuteCallbackHandlers时抛出了未处理的异常,直接导致应用崩溃。
解决方案
1. 注册全局未处理异常捕获(兜底方案)
在应用启动时添加全局异常处理,确保所有逃脱的异常都能被捕获记录,避免直接崩溃:
// 捕获AppDomain级别的未处理异常 AppDomain.CurrentDomain.UnhandledException += (sender, args) => { var exception = args.ExceptionObject as Exception; _logger.LogCritical(exception, "全局未处理异常"); // 执行必要的资源清理后优雅退出 }; // 捕获未被观察的Task异常 TaskScheduler.UnobservedTaskException += (sender, args) => { _logger.LogCritical(args.Exception, "未观察到的Task异常"); args.SetObserved(); // 标记异常已处理,阻止应用崩溃 };
2. 优化业务代码的异常处理逻辑
(1)给Polly传递CancellationToken
确保Parallel.ForAsync的取消令牌传递给Polly的ExecuteAsync方法,让取消信号能正确传递到底层操作:
await Parallel.ForAsync(0, totalBatchesCount, parallelOptions, async (i, cancellationToken) => { try { // 计算startId、endId逻辑... await retryPolicy.ExecuteAsync(async ct => { var batch = await _userRepository.GetBatchAsync(startId, endId, ct); // 处理批次数据逻辑... }, cancellationToken); // 传递取消令牌 } catch (Exception ex) { _logger.LogError(ex, $"批次{i}处理失败"); // 根据业务需求决定是否终止或继续 } });
(2)简化仓储层异常处理
仓储层无需额外捕获AggregateException,EF Core异步操作抛出的是具体异常(如MySqlException),直接让异常向上抛出由Polly处理即可:
public async Task<List<User>> GetBatchAsync(int startId, int endId, CancellationToken cancellationToken) { await using var context = await _contextFactory.CreateDbContextAsync(cancellationToken); return await context.Users .Where(u => u.UserId >= startId && u.UserId < endId) .Include(u => u.Invoice) .AsNoTracking() .ToListAsync(cancellationToken); }
3. 升级MySQL客户端包
这个底层异常可能是MySQL.Data包的已知问题,升级到最新稳定版本(或替换为MySqlConnector)可以修复这类超时回调的异常问题。
4. 替换Parallel.ForAsync为Task.WhenAll(可选)
如果Parallel.ForAsync的异常处理不够灵活,可以改用Task.WhenAll实现并行处理,异常控制更直观:
var tasks = Enumerable.Range(0, totalBatchesCount) .Select(async i => { try { // 计算startId、endId逻辑... await retryPolicy.ExecuteAsync(async ct => { var batch = await _userRepository.GetBatchAsync(startId, endId, ct); // 处理批次数据逻辑... }, cancellationToken); } catch (Exception ex) { _logger.LogError(ex, $"批次{i}处理失败"); } }); await Task.WhenAll(tasks);
内容的提问来源于stack exchange,提问作者Maciejosas

