如何用Vega实现按分组字段分面的词云?
实现Vega分面网格词云 + 引用Python变量方案
一、实现分组网格词云
Vega本身没有原生分面功能,但可以通过Repeat机制 + 数据过滤实现网格布局的分组词云,步骤如下:
1. 加载GitHub数据
import requests import json # 替换为你的GitHub数据URL data_url = "https://raw.githubusercontent.com/your-repo/your-data.json" raw_data = requests.get(data_url).json() # 提取所有分组(用于后续Repeat配置) all_groups = list({item["group"] for item in raw_data})
2. 构建分面Vega Spec
核心思路是用repeat定义网格的行/列分组,每个子视图通过Signal过滤对应分组的数据,生成独立词云:
# 基础词云Spec模板 base_wordcloud_spec = { "$schema": "https://vega.github.io/schema/vega/v5.json", "width": 250, "height": 250, "padding": 10, "signals": [{"name": "target_group", "value": ""}], "data": [ { "name": "filtered_data", "values": raw_data, "transform": [{"type": "filter", "expr": "datum.group == target_group"}] }, { "name": "wordcloud_data", "source": "filtered_data", "transform": [ # 统计词频(保留3个字符以上的词) {"type": "countpattern", "field": "text", "case": "lower", "pattern": "[\\w']{3,}", "stopwords": []}, # 随机旋转角度 {"type": "formula", "as": "angle", "expr": "random() * 6.283185307179586"}, # 词重(长词更粗) {"type": "formula", "as": "weight", "expr": "if(datum.text.length > 6, 700, 400)"}, # 生成词云布局 {"type": "wordcloud", "size": [{"signal": "width"}, {"signal": "height"}], "text": "text", "rotate": "angle", "fontSize": "count", "fontWeight": "weight", "font": "Arial", "padding": 3} ] } ], "marks": [ { "type": "text", "from": {"data": "wordcloud_data"}, "encode": { "enter": { "text": {"field": "text"}, "align": {"value": "center"}, "baseline": {"value": "alphabetic"}, "fill": {"value": "#2c3e50"} }, "update": { "x": {"field": "x"}, "y": {"field": "y"}, "rotate": {"field": "rotate"}, "fontSize": {"field": "fontSize"}, "fontWeight": {"field": "fontWeight"}, "fillOpacity": {"value": 1} }, "hover": {"fillOpacity": {"value": 0.6}} } } ] } # 构建分面网格Spec(按行排列分组,可改为column或同时row+column) faceted_spec = { "$schema": "https://vega.github.io/schema/vega/v5.json", "repeat": {"row": all_groups}, # 改为"column": all_groups则按列排列 "spec": base_wordcloud_spec, "resolve": {"signal": {"target_group": {"repeat": "row"}}} # 绑定Repeat变量到过滤Signal }
3. 在Colab中渲染
用你熟悉的两种方式都可以:
方式一:Altair Vega
from altair import vega vega.renderers.enable('colab') vega.Vega(faceted_spec)
方式二:Panel
import panel as pn from vega import Vega pn.extension('vega') pn.pane.Vega(faceted_spec)
二、在Vega Spec中引用Python变量
直接修改Spec的字典结构即可,无需复杂模板:
# 示例:动态设置词云尺寸、字体 custom_width = 300 custom_height = 300 custom_font = "Helvetica Neue" # 修改基础Spec的参数 base_wordcloud_spec["width"] = custom_width base_wordcloud_spec["height"] = custom_height base_wordcloud_spec["data"][1]["transform"][-1]["font"] = custom_font base_wordcloud_spec["data"][1]["transform"][-1]["size"] = [custom_width, custom_height] # 重新构建分面Spec(或直接修改已有的faceted_spec) faceted_spec["spec"] = base_wordcloud_spec
如果需要更复杂的动态逻辑(比如动态选择分组),可以将Spec转为字符串后用Python格式化,再转回JSON:
import json # 将Spec转为字符串 spec_str = json.dumps(base_wordcloud_spec) # 替换变量 spec_str = spec_str.replace('"font": "Arial"', f'"font": "{custom_font}"') # 转回JSON对象 updated_spec = json.loads(spec_str)
内容的提问来源于stack exchange,提问作者sisyphos
相关产品推荐
相关产品推荐

