BigQuery Storage API对clustered table的row restriction是否支持成本削减?
关于BigQuery Storage API对聚簇表的row restriction成本削减支持问题
我使用Scala调用BigQuery Storage API的Java客户端,读取一个包含4个聚簇字段(f1、f2、f3、f4)的聚簇表,创建该表的命令如下:
bq mk [...] --table --clustering_fields f1,f2,f3,f4 mytable mytableschema.json
表结构如下:
[ { "name": "f1", "type": "STRING", "mode": "REQUIRED" }, { "name": "f2", "type": "STRING", "mode": "NULLABLE" }, { "name": "f3", "type": "STRING", "mode": "REQUIRED" }, { "name": "f4", "type": "STRING", "mode": "REQUIRED" }, // other fields... ]
执行常规SQL查询时:
SELECT * FROM dataset.mytable WHERE f1 IN ('a', 'b') AND f2 IS NULL AND f3 = 'x'
系统会正确统计仅处理符合(a, null, x)和(b, null, x)聚簇的数据量。
但使用Storage API设置相同的row restriction读取时,账单显示按全表大小计费,API预估也显示全表将被计费。API核心使用逻辑如下:
val options = TableReadOptions .newBuilder() .setRowRestriction("f1 IN ('a', 'b') AND f2 IS NULL AND f3 = 'x'") .build() val readSessionBuilder = ReadSession .newBuilder() .setTable(tableName) .setDataFormat(DataFormat.AVRO) .setReadOptions(options) val readSessionRequestBuilder = CreateReadSessionRequest .newBuilder() .setParent(ProjectName) .setReadSession(readSessionBuilder) .setMaxStreamCount(1) val session = client.createReadSession(readSessionRequestBuilder.build()) // ... read from session.getStreamsList
请问BigQuery Storage API是否支持通过row restriction对聚簇表实现成本削减?我未找到相关说明文档。
内容的提问来源于stack exchange,提问作者Giovanni Caporaletti
相关产品推荐
相关产品推荐

