You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Vega中如何实现数据点聚合?能否像Vega-Lite一样在Encoding内聚合?

在Vega中实现聚合的正确方式

Vega和Vega-Lite的设计逻辑不同,Vega不支持像Vega-Lite那样在Encoding中直接指定聚合,你必须提前在transform阶段完成数据聚合,或者通过标记的信号逻辑处理(但前者是更规范、易维护的方案)。

你的尝试代码"x": {"signal": "scale(myscale, sum(datum['no of rooms']))"有问题——datum指向的是单条数据记录,sum()在这里只能计算当前单条记录的数值(结果就是它本身),根本做不了分组聚合。

下面给你两种可行的实现方案:

方案一:用Transform做聚合(推荐)

这是最符合Vega设计思路的方式,先通过aggregate转换按id分组计算平均房间数,再用聚合后的字段做编码:

{
  "data": {
    "name": "thedata",
    "values": [
      {"id": "house1", "no of rooms": 2},
      {"id": "house1", "no of rooms": 3},
      {"id": "house2", "no of rooms": 4},
      {"id": "house3", "no of rooms": 8}
    ]
  },
  "transform": [
    {"calculate": "datum['no of rooms'] <= 5 ? 'small' : 'large'", "as": "myType"},
    // 新增聚合转换:按id分组,计算no of rooms的平均值
    {
      "aggregate": [{"op": "mean", "field": "no of rooms", "as": "mean_rooms"}],
      "groupby": ["id", "myType"]
    }
  ],
  "marks": [
    {
      "type": "symbol",
      "from": {"data": "thedata"},
      "encode": {
        "update": {
          "y": {"scale": "yScale", "field": "id"},
          "x": {"scale": "xScale", "field": "mean_rooms"},
          "size": {"value": 200},
          "fill": {"value": "#4c78a8"}
        }
      }
    }
  ],
  "scales": [
    {
      "name": "yScale",
      "type": "band",
      "domain": {"data": "thedata", "field": "id"},
      "range": "height"
    },
    {
      "name": "xScale",
      "type": "linear",
      "domain": {"data": "thedata", "field": "mean_rooms"},
      "range": "width",
      "nice": true
    }
  ],
  "axes": [
    {"orient": "left", "scale": "yScale"},
    {"orient": "bottom", "scale": "xScale", "title": "Mean Number of Rooms"}
  ]
}

关键说明:

  • 新增的aggregate转换会按id和myType分组,计算出每个分组的平均房间数,存到mean_rooms字段里
  • 后续编码直接使用聚合后的mean_rooms字段,和Vega-Lite的效果完全一致

方案二:通过信号实现聚合(不推荐)

如果一定要绕开Transform,也可以通过信号结合data函数分组计算,但这种写法可读性差,维护成本高:

// 仅展示核心部分,完整代码需要补充scales、axes等
"marks": [
  {
    "type": "symbol",
    "from": {"data": "thedata", "groupby": ["id"]},
    "encode": {
      "update": {
        "y": {"scale": "yScale", "field": "id"},
        "x": {
          "signal": "scale('xScale', mean(data('thedata', {filter: datum.id === d.id}), 'no of rooms'))"
        },
        "size": {"value": 200}
      }
    }
  }
]

关键说明:

  • 使用groupby对标记按id分组
  • 通过data()函数过滤出当前分组的所有数据,再用mean()计算平均值
  • 这种方式不如Transform直观,数据量大时性能也会受影响,不建议使用

总结:优先用Transform做聚合,这是Vega的标准用法,逻辑清晰且易于维护。

内容的提问来源于stack exchange,提问作者Tom Martens

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 15:12:45