咨询:如何用Bokeh实现类似Pandas的日期降采样生成柱状图
Bokeh原生降采样支持说明
Great question! Let's break this down clearly:
核心结论
Bokeh本身并不原生提供像Pandas那样直接的resample或groupby+Grouper式的降采样API——毕竟Bokeh的核心定位是交互式可视化库,而非数据处理工具。不过这完全不影响你实现需求:你可以先用Pandas完成降采样/分组的预处理,再把处理后的数据传给Bokeh进行可视化;如果需要交互式切换采样粒度,也可以通过Bokeh的回调功能实现动态更新。
具体实现示例
假设你的原始DataFrame是df,我们先展示如何用Pandas预处理,再用Bokeh可视化:
1. 静态降采样(比如按月聚合)
先按月份降采样,同时按Col2的分类分组,计算Col3的总和:
import pandas as pd from bokeh.plotting import figure, show from bokeh.models import ColumnDataSource from bokeh.palettes import Category10_3 # 模拟原始DataFrame df = pd.DataFrame({ "Col1": pd.date_range(start="2023-01-01", periods=365, freq="D"), "Col2": pd.np.random.choice(["Cat", "Dog", "Bird"], size=365), "Col3": pd.np.random.randint(1, 10, size=365) }) # Pandas降采样:按月+分类分组,求和 resampled_df = df.groupby([pd.Grouper(key="Col1", freq="M"), "Col2"])["Col3"].sum().reset_index() # 转换为Bokeh的ColumnDataSource source = ColumnDataSource(resampled_df) # 绘制折线图 p = figure(x_axis_type="datetime", title="Monthly Count by Category") categories = ["Cat", "Dog", "Bird"] colors = Category10_3 for cat, color in zip(categories, colors): p.line( x="Col1", y="Col3", source=source, line_width=2, color=color, legend_label=cat, subset=source.data["Col2"] == cat ) p.legend.location = "top_left" show(p)
2. 交互式切换采样粒度(用Bokeh Server)
如果需要动态切换月/季/年粒度,你可以用Bokeh Server结合Python回调:
from bokeh.io import curdoc from bokeh.models import Dropdown # 初始化数据源 source = ColumnDataSource(data=dict(Col1=[], Col2=[], Col3=[])) categories = ["Cat", "Dog", "Bird"] colors = Category10_3 # 创建下拉菜单 freq_options = [("Month", "M"), ("Quarter", "Q"), ("Year", "Y")] dropdown = Dropdown(label="Sampling Frequency", menu=freq_options) # 定义回调函数 def update_data(attr, old, new): # 根据选择的粒度重新降采样 resampled = df.groupby([pd.Grouper(key="Col1", freq=new), "Col2"])["Col3"].sum().reset_index() source.data = ColumnDataSource(resampled).data dropdown.on_change("value", update_data) # 初始化图表 p = figure(x_axis_type="datetime", title="Dynamic Resampled Count by Category") for cat, color in zip(categories, colors): p.line(x="Col1", y="Col3", source=source, line_width=2, color=color, legend_label=cat) # 添加组件到文档 curdoc().add_root(p) curdoc().add_root(dropdown)
运行这个脚本需要用bokeh serve --show your_script.py命令启动Bokeh Server。
补充:Holoviews结合的便捷方式
既然你的DataFrame是被Holoviews封装的,其实可以直接利用Holoviews对Pandas resample的支持,再导出为Bokeh图表,这样更简洁:
import holoviews as hv hv.extension("bokeh") # 创建Holoviews的Curve元素,按分类分组 curve = hv.Curve(df, kdims=["Col1"], vdims=["Col3", "Col2"]).groupby("Col2") # 应用降采样,切换粒度只需修改freq参数 resampled_curve = curve.opts(resample={"freq": "M"}).opts(width=800, height=400) # 渲染为Bokeh图表 hv.render(resampled_curve)
内容的提问来源于stack exchange,提问作者Tyler Russell
相关产品推荐
相关产品推荐

