基于Redis存储贵金属时序数据用于ML的可行性及数据查询问询
Absolutely, your use case is totally feasible with Redis—let me walk you through how to make this work effectively, including handling date-specific and range queries, and supporting your ML workload.
1. Feasibility for High-Frequency Time-Series Data
Redis might be a key-value store at its core, but it’s got excellent support for time-series data, especially with the RedisTimeSeries module (built specifically for this kind of workload). Even without the module, you can use built-in data structures like Sorted Sets to handle your second-level gold/silver price data:
- RedisTimeSeries: This is the best fit for your needs. It’s optimized for high-throughput writes (perfect for second-level updates) and efficient range queries. It also supports automatic retention policies, so you can set it to keep exactly 52 weeks of data without manual cleanup.
- Sorted Sets: If you can’t use the module, you can map timestamps to prices using the timestamp as the
scoreand the price as themember. Writes are fast, and range queries are straightforward with sorted set commands.
Daily storage is easy—you can structure your keys to reflect dates (e.g., gold:price:20240704 for daily sorted sets, or a single gold:price time series with retention set to 52 weeks).
2. Querying Specific Dates & Date Ranges
Let’s break down how to pull the data you need, using both RedisTimeSeries and Sorted Sets:
Using RedisTimeSeries
- Create a time series with a 52-week retention (calculated as 527246060*1000 = 1814400000 milliseconds):
TS.CREATE gold:price RETENTION 1814400000 LABELS type precious_metal asset gold - Insert second-level data (the
*uses the current timestamp automatically):TS.ADD gold:price * 1902.35 - Query a specific date (e.g., July 4, 2024—convert the start/end of the day to milliseconds):
TS.RANGE gold:price 1688409600000 1688495999999 - Query a date range (e.g., Feb 1 to March 31, 2024):
TS.RANGE gold:price 1706736000000 1711823999999
You can also use TS.MRANGE if you need to pull data for multiple assets (gold and silver) in one query.
Using Sorted Sets (Without Modules)
- Daily key structure: Use keys like
gold:price:20240704for each day’s data. Insert data with:ZADD gold:price:20240704 1688409600 1902.35 - Query a specific date:
ZRANGEBYSCORE gold:price:20240704 1688409600 1688495999 WITHSCORES - Query a date range: You’ll need to iterate over all daily keys in the range (e.g.,
gold:price:20240201togold:price:20240331) and runZRANGEBYSCOREon each. To make this easier, you can maintain a separate sorted set that tracks all daily keys by their timestamp, so you can quickly fetch the keys in your date range first.
3. Supporting Your Machine Learning Workload
For ML, you’ll need to export the 52 weeks of data into a format your ML tools (like Python’s pandas/scikit-learn) can work with:
- With RedisTimeSeries, use
TS.RANGEorTS.MRANGEto pull all the data in one go, then parse the results into a DataFrame. - With Sorted Sets, batch-fetch the data from all relevant daily keys, combine the results, and convert to a structured format.
- If memory is a concern (52 weeks of second-level data is about 3.7 million points per asset), you can periodically archive older data to a disk-based store (like CSV files) while keeping the most recent data in Redis for low-latency access. Redis’s RDB/AOF persistence also ensures your data is safe between restarts.
Final Notes
Redis is more than capable of handling your high-frequency writes, date-based queries, and serving as a data source for ML. The RedisTimeSeries module will simplify your workflow significantly, but even without it, Sorted Sets are a solid alternative.
内容的提问来源于stack exchange,提问作者Nespl NS3

