You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kusto查询truncationmaxsize限制未生效返回数据超限问题

问题产生原因
  • 统计口径不匹配:estimate_data_size(*)返回的是数据在Kusto引擎内存中的未压缩存储大小,而truncationmaxsize的截断判定依据是结果集序列化后准备传输给客户端的压缩后数据大小,二者计算逻辑完全不同,你统计得到的10.7MB是内存未压缩值,不能作为参数是否生效的判断标准。
  • 参数生效时机不符:truncationmaxsize的截断动作发生在整个查询逻辑全部执行完成、即将向客户端返回结果的阶段,不会介入查询内部的计算流程。你在查询末尾追加的大小统计语句属于查询逻辑的一部分,会在截断校验前执行完成,自然可以得到超过你设定阈值的统计结果,这不属于参数失效。
  • 自定义参数值被服务端覆盖:Application Insights底层依赖的Azure Monitor服务对truncationmaxsize设有默认最小阈值(10MB,即10485760字节),如果用户设置的值低于该下限,服务端会静默忽略自定义值,使用默认阈值执行。你设置的8000000字节约为7.63MB,低于最小阈值,因此自定义配置根本没有生效。
  • 补充:truncationmaxrecords参数没有类似的最小阈值限制,所以可以正常按照你设置的值生效。
修复方法
  • 调整参数设置:如果要使用truncationmaxsize做保护,设置值不要低于10MB的服务端下限。
  • 不要依赖该参数做精确的结果大小控制,在查询逻辑层面提前裁剪数据,从源头降低结果集大小:
    • 移除查询返回中不需要的冗余列,仅保留业务必需字段
    • 根据单条数据的平均大小,调整top算子的返回行数,不要固定返回1000条
    • 对message这类长文本字段,使用substring()截取业务需要的长度,避免返回全量超长文本
  • 如果需要严格控制返回结果的内存大小在8MB以内,可以通过查询逻辑实现精确截断,参考写法:
set truncationmaxsize=10485760;
union (traces | extend details = dynamic(null)), (exceptions | project timestamp, operation_Id, name = iff(isnotempty(innermostType), innermostType, outerType), message = iff(isnotempty(innermostMessage), innermostMessage, outerMessage), details) 
| where timestamp >= datetime(2022-07-06T22:28:43.539650600) 
| as main 
| project operation_Id, operation_Name, Timestamp = timestamp, SeverityLevel = severityLevel, Name = name, Message = message 
| where Message contains 'Job failed. Error:' 
| summarize arg_max(Timestamp, *) by operation_Id 
| join kind=leftouter (main | project operation_Id, except_timestamp=timestamp, except_severityLevel = severityLevel, except_Message = message | where (isempty(except_severityLevel) and except_Message !contains "retry") | summarize arg_max(except_timestamp, *) by operation_Id) on operation_Id 
| join kind=leftouter (main | project operation_Id, dt_message = message, dt_timestamp = timestamp | where dt_message contains 'Type Data' or dt_message contains 'Type: Data' or dt_message contains 'Execute data' | summarize arg_max(dt_timestamp, *) by operation_Id | extend data_type = extract(@"(?:Type|Type:|Execute)\s*((?i)data\s*\w+)", 1, dt_message) | project operation_Id, dt_timestamp, dt_message, failed_on = data_type) on operation_Id 
| join kind=leftouter (main | project operation_Id, activity_number_message = message, anm_timestamp = timestamp | where activity_number_message contains 'Processing activity' | summarize arg_max(anm_timestamp, *) by operation_Id | project operation_Id, anm_timestamp, activity_number_message) on operation_Id 
| sort by Timestamp asc
| extend row_size = estimate_data_size(*)
| serialize cumulative_size = row_cumsum(row_size)
| where cumulative_size <= 8000000
| summarize Total=sum(row_size)
  • 校验参数实际生效状态:在查询开头添加set query_stats=true;,执行后查看查询返回的统计信息中的DataSetSize字段,该值才是truncationmaxsize截断判定使用的实际序列化大小,不要用estimate_data_size(*)的结果作为判定依据。

原始问题中使用的查询语句如下:

set truncationmaxsize=8000000; union (traces | extend details = dynamic(null)), (exceptions | project timestamp, operation_Id, name = iff(isnotempty(innermostType), innermostType, outerType), message = iff(isnotempty(innermostMessage), innermostMessage, outerMessage), details) | where timestamp >= datetime(2022-07-06T22:28:43.539650600) | as main | project operation_Id, operation_Name, Timestamp = timestamp, SeverityLevel = severityLevel, Name = name, Message = message | where Message contains 'Job failed. Error:' | summarize arg_max(Timestamp, *) by operation_Id | join kind=leftouter (main | project operation_Id, except_timestamp=timestamp, except_severityLevel = severityLevel, except_Message = message | where (isempty(except_severityLevel) and except_Message !contains "retry") | summarize arg_max(except_timestamp, *) by operation_Id) on operation_Id | join kind=leftouter (main | project operation_Id, dt_message = message, dt_timestamp = timestamp | where dt_message contains 'Type Data' or dt_message contains 'Type: Data' or dt_message contains 'Execute data' | summarize arg_max(dt_timestamp, *) by operation_Id | extend data_type = extract(@"(?:Type|Type:|Execute)\s*((?i)data\s*\w+)", 1, dt_message) | project operation_Id, dt_timestamp, dt_message, failed_on = data_type) on operation_Id | join kind=leftouter (main | project operation_Id, activity_number_message = message, anm_timestamp = timestamp | where activity_number_message contains 'Processing activity' | summarize arg_max(anm_timestamp, *) by operation_Id | project operation_Id, anm_timestamp, activity_number_message) on operation_Id | top 1000 by Timestamp asc | summarize Total=sum(estimate_data_size(*)) 

内容的提问来源于stack exchange,提问作者beraiyan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 20:55:01