You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

启用CBO与自动统计收集时,MSCK REPAIR是否生成表统计信息?

Hive MSCK REPAIR vs. ANALYZE for CBO Statistics

Great question—let’s break this down clearly so you know exactly what to expect.

Core Behavior Explanation:

  • MSCK REPAIR TABLE <table> exists solely to sync Hive’s metastore metadata with the actual partition directories on HDFS. It adds new partitions or removes stale ones from the metastore, but it will NOT generate or update the table/partition statistics needed for CBO—even when hive.cbo.enable=true and hive.stats.autogather=true are enabled.

  • The hive.stats.autogather=true parameter only triggers automatic statistics collection during data-writing operations like INSERT, LOAD DATA, or CREATE TABLE AS SELECT. Since MSCK is just a metadata synchronization tool (not a data ingestion/write action), it doesn’t trigger this auto-gather logic.

  • To get the accurate statistics the Cost-Based Optimizer (CBO) relies on (like row counts, data sizes, and column distribution), you must explicitly run the ANALYZE command:

    • For full table statistics: ANALYZE TABLE <tablename> COMPUTE STATISTICS
    • For new partitions synced via MSCK: ANALYZE TABLE <tablename> PARTITION (<partition_col>=<value>) COMPUTE STATISTICS

This ensures the CBO has the data it needs to build optimized query execution plans.

内容的提问来源于stack exchange,提问作者Yaniv Irony

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 21:48:13