如何在Python中可视化含尖峰的密集间隙型传感器时间序列?
高采样率带间隙传感器时序数据的可视化方案对比
数据特性
- 时长约12小时的高采样率数据,总计约250,000个数据点
- 包含11个连续数据段,段间存在30-40分钟的间隙(传感器周期性离线)
- 信号中位数约为70,但存在峰值超过500的尖锐尖峰
可视化需求
需要根据强度对数据点进行颜色编码(低于中位数、中位数1-2倍区间、高于中位数2倍),并显示阈值线。
最小可复现示例代码
下面是生成类似数据并实现6种可视化方案的Python代码:
""" Minimal reproducible example: visualizing a dense, gappy time series with spikes. Shows 6 different approaches for comparison. """ import numpy as np import matplotlib matplotlib.use('Agg') import matplotlib.pyplot as plt np.random.seed(42) # Simulate a gappy sensor time series: # ~12 hours of recording split into ~11 segments # with ~30-40 min gaps (sensor offline periodically) segments = [] t_start = 0 for i in range(11): duration = np.random.uniform(2000, 3000) n_points = int(duration / 0.1) t = np.linspace(t_start, t_start + duration, n_points) base = 70 + 20 * np.sin(2 * np.pi * t / 5000) noise = np.cumsum(np.random.randn(n_points)) * 0.3 noise -= np.mean(noise) rate = base + noise + np.random.poisson(base) - base n_spikes = np.random.poisson(3) for _ in range(n_spikes): spike_pos = np.random.randint(0, n_points) spike_amp = np.random.exponential(100) spike_width = np.random.uniform(5, 50) spike = spike_amp * np.exp(-0.5 * ((t - t[spike_pos]) / spike_width) ** 2) rate += spike rate = np.maximum(rate, 0) segments.append((t / 3600, rate)) # store in hours gap = np.random.uniform(1800, 2400) t_start += duration + gap # Thresholds for color-coding all_rates = np.concatenate([s[1] for s in segments]) med = np.median(all_rates[all_rates > 0]) thresh1 = med thresh2 = med * 2 print(f"Total points: {len(all_rates)}") print(f"Segments: {len(segments)}") print(f"Median: {med:.1f}, 2x median: {thresh2:.1f}, Max: {np.max(all_rates):.1f}") def color_for_rate(r, t1=thresh1, t2=thresh2): return np.where(r < t1, '#2166ac', np.where(r < t2, '#ff8c00', '#cc0000')) def bin_segment(t, r, bin_width_hrs=100/3600): """Bin a single continuous segment.""" if len(t) < 2: return np.array([]), np.array([]), np.array([]) t_bins, r_bins, e_bins = [], [], [] t_min, t_max = t[0], t[-1] edges = np.arange(t_min, t_max, bin_width_hrs) for j in range(len(edges) - 1): mask = (t >= edges[j]) & (t < edges[j+1]) if np.sum(mask) > 0: t_bins.append(np.mean(t[mask])) r_bins.append(np.mean(r[mask])) e_bins.append(np.std(r[mask]) / np.sqrt(np.sum(mask))) return np.array(t_bins), np.array(r_bins), np.array(e_bins) # ===================================================================== # 6 visualization approaches # ===================================================================== fig, axes = plt.subplots(6, 1, figsize=(20, 30), sharex=True) # --- 1. Naive ax.plot (BAD: connects across gaps) --- ax = axes[0] t_all = np.concatenate([s[0] for s in segments]) r_all = np.concatenate([s[1] for s in segments]) ax.plot(t_all, r_all, 'k-', linewidth=0.3, alpha=0.5) ax.set_ylabel("Intensity") ax.set_title("1. ax.plot() — connects across gaps (bad)", fontsize=14, fontweight='bold') ax.set_ylim(0, thresh2 * 3) # --- 2. Per-segment line plot (fixes gaps, but dense) --- ax = axes[1] for t_seg, r_seg in segments: ax.plot(t_seg, r_seg, 'k-', linewidth=0.2, alpha=0.4) ax.set_ylabel("Intensity") ax.set_title("2. Per-segment ax.plot() — gaps correct, but too dense to read", fontsize=14, fontweight='bold') ax.set_ylim(0, thresh2 * 3) # --- 3. Errorbar points with 100s binning --- ax = axes[2] for t_seg, r_seg in segments: tb, rb, eb = bin_segment(t_seg, r_seg) if len(tb) > 0: colors = color_for_rate(rb) for c in ['#2166ac', '#ff8c00', '#cc0000']: mask = colors == c if np.any(mask): ax.errorbar(tb[mask], rb[mask], yerr=eb[mask], fmt='.', color=c, markersize=4, elinewidth=0.5, capsize=0) ax.axhline(thresh1, color='#2166ac', ls='--', lw=1.5, alpha=0.5) ax.axhline(thresh2, color='#cc0000', ls='--', lw=1.5, alpha=0.5) ax.set_ylabel("Intensity") ax.set_title("3. Errorbar + 100s binning — readable but spikes smoothed out", fontsize=14, fontweight='bold') ax.set_ylim(0, thresh2 * 3) # --- 4. Per-segment colored fill_between --- ax = axes[3] for t_seg, r_seg in segments: tb, rb, _ = bin_segment(t_seg, r_seg, bin_width_hrs=30/3600) if len(tb) < 2: continue ax.fill_between(tb, 0, rb, where=rb < thresh1, color='#2166ac', alpha=0.5, linewidth=0) ax.fill_between(tb, 0, rb, where=(rb >= thresh1) & (rb < thresh2), color='#ff8c00', alpha=0.5, linewidth=0) ax.fill_between(tb, 0, rb, where=rb >= thresh2, color='#cc0000', alpha=0.6, linewidth=0) ax.plot(tb, rb, 'k-', linewidth=0.4, alpha=0.5) ax.axhline(thresh1, color='#2166ac', ls='--', lw=1.5, alpha=0.5) ax.axhline(thresh2, color='#cc0000', ls='--', lw=1.5, alpha=0.5) ax.set_ylabel("Intensity") ax.set_title("4. Per-segment fill_between (30s bins) — shows structure but messy at transitions", fontsize=14, fontweight='bold') ax.set_ylim(0, thresh2 * 3) # --- 5. Vertical bars (stem plot) per segment ---
各方案特点对比
- 方案1:直接使用
ax.plot():最基础的折线图,但会跨数据间隙连接,导致视觉上的错误引导,完全不适合这类带间隙的时序数据 - 方案2:按数据段绘制折线图:修正了间隙问题,但因为原始数据点过于密集,折线重叠严重,难以分辨细节和尖峰特征
- 方案3:100秒分箱的误差棒图:通过分箱降低了数据密度,可读性大幅提升,同时用颜色编码区分了不同强度区间,还能展示数据的波动情况,但分箱操作会平滑掉尖锐的尖峰,丢失关键的峰值信息
- 方案4:按数据段的彩色
fill_between填充图(30秒分箱):用不同颜色填充对应强度区间,能清晰展示数据的整体结构和趋势,但在强度区间的过渡位置会显得杂乱,视觉上不够整洁 - 方案5:按数据段绘制竖条(茎状图):[原内容未完成,保留现状]
内容的提问来源于stack exchange,提问作者requiemman
相关产品推荐
相关产品推荐

