You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Application Insights中使用采样时,能否保留被排除项的相关条目?

在Application Insights中使用采样时,能否保留被排除项的相关条目?

这个需求我太懂了——既要靠采样把App Insights的日志量压在限额内,又不想错过异常,还得保留异常关联的完整调用链,不然排查问题时缺了关键痕迹可太闹心了。答案是完全可以实现,我给你说个最靠谱的落地方案:

核心思路:跟踪异常关联的Operation ID,让同ID的所有条目跳过采样

Application Insights的默认采样是逐个处理Telemetry条目,但我们可以自定义一个ITelemetryProcessor,通过维护一个“需要保留的Operation ID列表”,让关联的痕迹、请求等都跟着异常一起绕过采样逻辑。具体步骤如下:

  1. 实现自定义采样处理器
    我们需要一个线程安全的缓存来记录哪些Operation ID因为关联异常需要保留,同时给缓存加过期时间避免内存泄漏。处理器的逻辑是:

    • 遇到ExceptionTelemetry直接保留,把它的Operation ID加入缓存
    • 后续遇到任何和缓存中ID匹配的Telemetry(比如Trace、Request),也直接保留
    • 其他条目走正常的采样逻辑(比如你设置的50%概率)

    给你贴个可直接用的示例代码:

    using Microsoft.ApplicationInsights.Channel;
    using Microsoft.ApplicationInsights.Extensibility;
    using System.Collections.Concurrent;
    
    public class ExceptionLinkedSamplingProcessor : ITelemetryProcessor
    {
        private readonly ITelemetryProcessor _nextProcessor;
        private readonly ConcurrentDictionary<string, DateTimeOffset> _reservedOperationIds = new();
        private readonly double _targetSamplingRate;
        private readonly TimeSpan _cacheExpiry = TimeSpan.FromMinutes(5); // 可根据业务调整过期时间
    
        public ExceptionLinkedSamplingProcessor(ITelemetryProcessor nextProcessor, double samplingPercentage)
        {
            _nextProcessor = nextProcessor;
            _targetSamplingRate = samplingPercentage;
    
            // 启动定时任务清理过期的Operation ID,防止内存溢出
            _ = Task.Run(async () =>
            {
                while (true)
                {
                    await Task.Delay(TimeSpan.FromMinutes(1));
                    CleanupExpiredIds();
                }
            });
        }
    
        public void Process(ITelemetry telemetryItem)
        {
            // 先顺手清一波过期的ID,双重保险
            CleanupExpiredIds();
    
            var currentOperationId = telemetryItem.Context.Operation.Id;
            bool shouldKeepItem = false;
    
            // 1. 如果是异常,强制保留并记录Operation ID
            if (telemetryItem is ExceptionTelemetry)
            {
                shouldKeepItem = true;
                _reservedOperationIds.AddOrUpdate(currentOperationId, DateTimeOffset.UtcNow, (_, _) => DateTimeOffset.UtcNow);
            }
            // 2. 如果当前条目属于已标记要保留的调用链,也强制保留
            else if (_reservedOperationIds.ContainsKey(currentOperationId))
            {
                shouldKeepItem = true;
            }
            // 3. 其他情况走正常采样逻辑
            else
            {
                var random = new Random();
                shouldKeepItem = random.NextDouble() * 100 <= _targetSamplingRate;
            }
    
            if (shouldKeepItem)
            {
                _nextProcessor.Process(telemetryItem);
            }
        }
    
        private void CleanupExpiredIds()
        {
            var now = DateTimeOffset.UtcNow;
            var expiredIds = _reservedOperationIds
                .Where(kv => now - kv.Value > _cacheExpiry)
                .Select(kv => kv.Key)
                .ToList();
    
            foreach (var id in expiredIds)
            {
                _reservedOperationIds.TryRemove(id, out _);
            }
        }
    }
    
  2. 注册自定义处理器并关闭默认采样
    接下来在你的Program.cs里,把这个自定义处理器注册进去,记得关掉Application Insights的默认自适应采样,避免逻辑冲突:

    var services = new ServiceCollection();
    
    // 配置Application Insights,关闭默认采样
    services.AddApplicationInsightsTelemetry(options =>
    {
        options.ConnectionString = appInsightConnString;
        options.EnableAdaptiveSampling = false; // 必须关掉,不然和自定义采样冲突
        options.EnableQuickPulseMetricStream = true; // 按需保留
    })
    // 注册我们的自定义采样处理器
    .AddTelemetryProcessor(sp =>
    {
        var nextProcessor = sp.GetRequiredService<ITelemetryProcessor>();
        return new ExceptionLinkedSamplingProcessor(nextProcessor, samplingPercentage);
    });
    
    // 后续的服务构建和使用逻辑不变
    var serviceProvider = services.BuildServiceProvider();
    // ...
    

几个要注意的小细节

  • 线程安全:一定要用ConcurrentDictionary这种线程安全的集合存Operation ID,不然多线程处理日志时容易出问题
  • 内存占用:缓存过期时间别设太长,比如5分钟足够覆盖大部分调用链的生命周期了,定时清理也别忘了
  • 采样逻辑适配:如果你想用自适应采样而不是固定比例,也可以把自定义处理器里的固定采样逻辑换成调用内置的自适应采样处理器,灵活调整就行

总的来说,这个方案能完美满足你的需求:既按比例采样大部分日志,又能完整保留所有异常以及它们关联的全部调用痕迹,排查问题时再也不会缺东少西啦。

备注:内容来源于stack exchange,提问作者Alex K.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.17 09:29:45