BigTable中column family名称是否会在每行中重复存储?
结论:BigTable 会在每行的列项中存储 column family 的标识,但可能通过预创建特性做编码优化;从空间优化的最佳实践角度,仍建议使用短名称来控制开销。
核心分析
官方文档的隐含逻辑
官方文档明确说明,行的本质是键值条目集合,每个列项的键由column family、column qualifier和时间戳组合而成,这意味着 column family 的信息会伴随每行的每个列项存储。row is essentially a collection of key/value entries, where the key is a combination of the column family, column qualifier and timestamp.
HBase 的参考类比
HBase 作为基于 BigTable 论文实现的系统,明确要求缩短 column family 和 qualifier 的名称——因为二者会在每行重复存储,短名称能减少读写的数据量。BigTable 作为原型系统,逻辑上大概率遵循相同的存储模式。The column family and column qualifier names are repeated for each row. Therefore, keep the names as short as possible to reduce the amount of data that HBase stores and reads.
预创建特性的优化可能性
不同于写入时动态创建的column qualifier,column family必须提前通过控制台、CLI 或 API 创建。BigTable 引擎可利用这一点做编码优化(比如用短ID替代完整名称存储,读写时再映射回原名),但官方文档并未明确这一优化的存在,因此不能依赖它来节省空间。
官方文档关于 Qualifier 的明确提示
官方文档仅明确指出 column qualifier 会在每行重复存储,建议将数据本身用作限定符来节省空间,但未提及 column family 的空间细节:
Treat column qualifiers as data. Since you have to store a column qualifier for every column, you can save space by naming the column with a value.
实践建议
- 稳妥起见,建议继续使用短名称作为 column family 的标识(如用
d表示 default、m表示 metadata),避免重复存储完整名称带来不必要的空间开销。 - 若后续官方文档明确说明 column family 无需重复存储完整名称,再考虑使用可读性更好的完整名称。
内容的提问来源于stack exchange,提问作者guruprasadbhat

