You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现滚动式数据流重采样的方案及工具咨询

Great question—handling rolling resampling for large time-series datasets while keeping memory usage and frontend data transfer in check is a common pain point when building Dash apps. Here are my top Python library recommendations tailored exactly to your use case:

1. Pandas + NumPy (Simple, Integrated with Dash)

Pandas is the go-to for time-series work, and paired with NumPy, it’s easy to implement dynamic rolling resampling for your use case. Since you need to maintain exactly 200 resampled points once you exceed the initial threshold, you can use uniform index sampling or aggregation bins:

import pandas as pd
import numpy as np

# Maintain a list of raw incoming data points (e.g., vehicle speeds)
raw_data = []

def update_resampled_data(new_speed):
    raw_data.append(new_speed)
    sample_count = len(raw_data)
    
    if sample_count <= 200:
        # Keep all points initially
        return raw_data.copy()
    else:
        # Generate 200 evenly spaced indices across the full raw dataset
        resample_indices = np.linspace(0, sample_count - 1, 200, dtype=int)
        # Extract the resampled points
        return [raw_data[i] for i in resample_indices]

If your data includes timestamps, use Pandas resample to aggregate (mean/max/median) into 200 time bins instead—this preserves temporal context better than uniform index sampling.

2. SciPy Signal Resampling (Smooth, High-Quality)

For smoother resampling (especially useful for continuous data like vehicle speed), SciPy’s signal.resample uses FFT-based interpolation to downsample your dataset to exactly 200 points. It’s ideal if you want to preserve the overall trend without losing too much detail:

from scipy import signal
import numpy as np

raw_data = []

def update_resampled_data(new_speed):
    raw_data.append(new_speed)
    sample_count = len(raw_data)
    
    if sample_count <= 200:
        return np.array(raw_data)
    else:
        # Resample to exactly 200 points using FFT
        return signal.resample(raw_data, 200).tolist()

Note: This has a tiny bit more computational overhead than simple index sampling, but it’s negligible for updating a Dash chart in real time.

3. Dask (Out-of-Core Processing for Extreme Data Sizes)

If your dataset is so large it can’t fit into memory (millions of records), Dask is perfect. It mimics Pandas’ API but processes data in chunks, avoiding RAM overload. You can use Dask DataFrames to handle rolling resampling without loading all data at once:

import dask.dataframe as dd
import numpy as np

# Assume you're reading data in chunks from your database
chunk_iterator = your_database_query.read_chunks(chunk_size=1000)
raw_data = []

for chunk in chunk_iterator:
    raw_data.extend(chunk['speed'].tolist())
    if len(raw_data) > 200:
        # Convert to Dask Series for out-of-core resampling
        dask_series = dd.from_pandas(pd.Series(raw_data), npartitions=2)
        # Resample to 200 points (adjust aggregation as needed)
        resampled = dask_series.resample(f"{len(raw_data)//200}ms").mean().compute()
        raw_data = resampled.tolist()  # Reset to resampled points to keep memory low

Dask’s strength is handling datasets that would crash Pandas due to memory limits, while keeping your code familiar if you already know Pandas.

4. River (Formerly Creme) - Streaming-First Processing

If you’re dealing with a continuous stream of data (not just a large static dataset), River is built for online/streaming workflows. It lets you maintain a sliding window of raw data and dynamically resample to 200 points without storing the entire history:

from river import utils
import numpy as np

# Maintain a sliding window of raw data (adjust max size if needed)
window = utils.SlidingWindow(size=1000)  # Keep last 1000 points to resample from

def update_resampled_data(new_speed):
    window.append(new_speed)
    window_size = len(window)
    
    if window_size <= 200:
        return list(window)
    else:
        # Resample the current window to 200 points
        resample_indices = np.linspace(0, window_size - 1, 200, dtype=int)
        return [window[i] for i in resample_indices]

River is lightweight and designed for low-latency streaming, making it great if your data is coming in continuously (e.g., real-time vehicle telemetry).


Which Should You Choose?

  • Pandas + NumPy: Best for most cases—simple, easy to integrate with Dash, and sufficient for datasets that fit in memory.
  • SciPy: If you need smoother, higher-quality resampling for continuous data.
  • Dask: If your dataset is too large to fit in RAM (millions of records with high memory footprint).
  • River: If you’re working with a real-time stream and need minimal memory overhead.

All these libraries output a list/array of exactly 200 points, which you can directly pass to Dash’s dcc.Graph component—this keeps frontend data transfer tiny and avoids performance issues.

内容的提问来源于stack exchange,提问作者mfcss

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 14:47:48