如何将大量时间序列绘制成带量化密度的灰度渐变图?
多时间序列的灰度密度可视化方案
针对你提到的2000条时间序列可视化的痛点——直接df.plot()杂乱无章、仅看均值缺乏密度量化信息,这里分享几个亲测好用的专业实现方法,完美满足你用灰度渐变体现数值区间序列密度的需求:
方案1:区间分箱热力图(量化离散区间的序列数量)
这个方法把时间轴和数值轴都切成固定区间,统计每个网格内的序列点数量,用灰度深浅对应密度高低,搭配色条明确量化数值,完全避开alpha值的临时方案。
import pandas as pd import numpy as np import matplotlib.pyplot as plt from matplotlib.colors import LinearSegmentedColormap # 生成模拟数据(模拟2000条时间序列) ts = pd.Series(np.random.randn(1000), index=pd.date_range('1/1/2000', periods=1000)) df = pd.DataFrame(np.random.randn(1000, 2000), index=ts.index) # 1. 定义时间和数值的分箱区间 time_bins = pd.cut(df.index.astype(np.int64)//10**9, bins=50) # 把时间切成50个均匀区间 value_bins = pd.cut(df.stack(), bins=[-4, -2, -1, 0, 1, 2, 4]) # 自定义你需要的数值区间 # 2. 统计每个(time_bin, value_bin)的序列点数量 count_df = df.stack().groupby([time_bins, value_bins]).count().unstack().fillna(0) # 3. 自定义灰度渐变colormap(深灰对应高密度,浅灰对应低密度) gray_cmap = LinearSegmentedColormap.from_list('gray_grad', ['#f0f0f0', '#333333']) # 4. 绘制热力图 plt.figure(figsize=(12,6)) im = plt.imshow(count_df.T, aspect='auto', cmap=gray_cmap, extent=[df.index[0], df.index[-1], -4, 4]) plt.colorbar(im, label='序列点数量') plt.ylabel('数值区间') plt.xlabel('时间') plt.title('时间序列数值区间密度热力图') plt.xticks(rotation=45) plt.tight_layout() plt.show()
这个方案的核心优势是完全量化了每个区间的序列数量,色条直接给出密度参考,适合需要精确观察不同区间分布规律的场景。
方案2:连续核密度估计(KDE)的灰度填充
如果想要更平滑的连续密度展示,不需要手动定义分箱区间,可以用核密度估计计算整个时间-数值平面的密度分布,再用灰度渐变填充:
import seaborn as sns # 整理数据格式:时间转数值型Unix时间戳,方便KDE计算 time_values = df.index.astype(np.int64)//10**9 flat_time = np.repeat(time_values, df.shape[1]) flat_values = df.values.flatten() # 绘制连续密度图 plt.figure(figsize=(12,6)) # 用seaborn的kdeplot生成连续密度,cmap用灰度渐变,shade=True实现填充 sns.kdeplot(x=flat_time, y=flat_values, cmap=gray_cmap, shade=True, levels=10, cbar=True, cbar_kws={'label':'密度'}) # 把x轴转回日期格式,提升可读性 plt.xticks(ticks=plt.xticks()[0], labels=pd.to_datetime(plt.xticks()[0]*10**9).strftime('%Y-%m-%d'), rotation=45) plt.ylabel('数值') plt.xlabel('时间') plt.title('时间序列数值连续密度分布') plt.tight_layout() plt.show()
这个方案视觉上更平滑自然,适合观察整体的密度趋势,KDE会自动计算最优的密度分布,不需要人工干预分箱规则。
方案3:分层分位数灰度填充(类箱线图的连续版本)
如果你想保留类似箱线图的分位数统计意义,同时结合你之前的滚动均值需求,可以计算序列的滚动分位数,然后用fill_between分层填充灰度区间:
# 计算滚动分位数(窗口60,和你之前的滚动均值窗口保持一致) rolling_q0 = df.rolling(window=60, center=True).quantile(0.0) rolling_q25 = df.rolling(window=60, center=True).quantile(0.25) rolling_q50 = df.rolling(window=60, center=True).quantile(0.5) rolling_q75 = df.rolling(window=60, center=True).quantile(0.75) rolling_q100 = df.rolling(window=60, center=True).quantile(1.0) # 自定义灰度层级(从深到浅对应中间到两端的分位数区间) gray_levels = ['#222222', '#555555', '#888888', '#bbbbbb'] plt.figure(figsize=(12,6)) # 分层填充分位数区间 plt.fill_between(df.index, rolling_q0, rolling_q25, color=gray_levels[3], alpha=1) plt.fill_between(df.index, rolling_q25, rolling_q50, color=gray_levels[2], alpha=1) plt.fill_between(df.index, rolling_q50, rolling_q75, color=gray_levels[1], alpha=1) plt.fill_between(df.index, rolling_q75, rolling_q100, color=gray_levels[0], alpha=1) # 叠加滚动均值线,和你之前的需求衔接 rolling_mean = df.mean(1).rolling(window=60, center=True).mean() rolling_mean.plot(color='white', linewidth=2, label='滚动均值') plt.legend() plt.colorbar(plt.cm.ScalarMappable(cmap=gray_cmap), label='分位数区间(深=75-100%,浅=0-25%)') plt.ylabel('数值') plt.xlabel('时间') plt.title('滚动分位数灰度分层密度图') plt.tight_layout() plt.show()
这个方案完美结合了滚动统计和密度可视化,每个灰度层都对应明确的统计意义,非常适合观察数据分布的离散程度随时间的变化趋势。
内容的提问来源于stack exchange,提问作者JC_CL
相关产品推荐
相关产品推荐

