Haskell不规则数据序列管理与指定分辨率重采样实现方法
问题域说明
现有如下预定义的Haskell类型与辅助函数:
- 时间戳类型别名
type TimeStamp = Int
- 带时间戳的单点数据类型
DataPoint,已派生Show、Eq实例,同时实现Foldable类型类:
data DataPoint a = DataPoint {index :: TimeStamp, value :: a} deriving (Show, Eq) instance Foldable DataPoint where foldMap f (DataPoint _ y) = f y
- 时间序列类型
Series,实现Foldable类型类时长度计算直接调用后续的size函数:
data Series a = Series [DataPoint a] instance Foldable Series where foldMap f (Series xs) = foldMap (foldMap f) xs length = size
- 配套辅助函数:
-- 构造空时间序列 emptySeries :: Series a emptySeries = Series [] -- 从(时间戳,值)的列表构造时间序列 timeSeries :: [(TimeStamp, a)] -> Series a timeSeries xs = Series $ map (uncurry DataPoint) xs -- 计算时间序列长度 size :: Series a -> Int size (Series xs) = length xs
实现需求
需要实现一个重采样函数,满足以下要求:
- 输入不规则时间数据集、目标分辨率,输出对应分辨率的
timeSeries - 每个分辨率对应的时间点,存储该时间点之前最新的有效值
- 最新数据放在序列最前端,提升访问效率
- 示例:输入不规则数据
irrData = [(98,5), (96,4), (93,9)],分辨率为1时,输出序列为[(98,5), (97,4), (96,4), (95,9), (94,9), (93,9)]
额外优化要求:
- 同一函数需支持将已有
timeSeries重采样为其他分辨率 - 可正确处理任意顺序传入的输入数据,例如乱序输入
[(98,5), (101,2), (93,4)]也能输出正确结果
当前实现进展
注:当前给出的解决方案已可实现预期功能,后续会随迭代优化更新为最终版本,目前待优化点为代码精简,存在单行代码过长的问题。
当前代码如下:
resample :: Int -> [(Int, v)] -> [(Int, v)] resample _ [] = [] resample _ [x] = [x] resample r xs = foldr (\i acc -> (i, snd $ head (filter (\x -> (fst x) <= i) s)):acc) [] [(fst $ head s), (fst $ head s) - r .. (fst $ last s)] where s = sortBy (flip $ on compare fst) xs
内容的提问来源于stack exchange,提问作者Reid Johnson
相关产品推荐
相关产品推荐

