修复JavaScript时间序列数据补全代码:填充缺失0值记录
修复时间序列数据缺口补全函数
问题背景
API返回的小时粒度时间序列数据中,所有EventsStaging.userCount为0的记录被省略,需要在客户端补全这些缺失的条目,确保时间序列连续,缺失时段的userCount设为0。
输入数据示例:
[ { "EventsStaging.timestamp.hour":"2022-07-31T20:00:00.000", "EventsStaging.timestamp":"2022-07-31T20:00:00.000", "EventsStaging.userCount":"1" }, { "EventsStaging.timestamp.hour":"2022-07-31T21:00:00.000", "EventsStaging.timestamp":"2022-07-31T21:00:00.000", "EventsStaging.userCount":"1" }, { "EventsStaging.timestamp.hour":"2022-08-01T04:00:00.000", "EventsStaging.timestamp":"2022-08-01T04:00:00.000", "EventsStaging.userCount":"1" } ]
目标输出需补全2022-07-31T22:00到2022-08-01T03:00之间的条目,每个缺失条目的userCount为0,且两个时间字段值一致。
原代码的问题
- 循环逻辑完全错误:原循环的终止条件写反,无法生成中间缺失的时间点
- 未添加当前条目:reduce过程中只尝试补全缺失项,没有把当前遍历到的有效条目加入结果集
- 时间字段处理不全:每个条目需要同时设置
EventsStaging.timestamp.hour和EventsStaging.timestamp,原代码只处理了一个字段 - 依赖字段顺序不可靠:通过
Object.keys(series[0])[0]和[2]取字段名,一旦API返回字段顺序变化就会出错 - 初始状态处理不当:当结果集为空时,
prev是空对象,补全的条目会缺失必要字段
修复后的代码
const fillMissingEntries = (series, granularity) => { // 定义时间步长,小时粒度为3600秒 const step = granularity === 'hour' ? 60 * 60 * 1000 : 0; if (step === 0 || series.length === 0) return [...series]; // 明确指定字段名,避免依赖返回顺序 const timeHourKey = 'EventsStaging.timestamp.hour'; const timeKey = 'EventsStaging.timestamp'; const countKey = 'EventsStaging.userCount'; return series.reduce((acc, currentEntry) => { const currentTime = new Date(currentEntry[timeHourKey]).getTime(); let lastTime; if (acc.length === 0) { // 第一个条目直接加入 acc.push({...currentEntry}); lastTime = currentTime; } else { lastTime = new Date(acc[acc.length - 1][timeHourKey]).getTime(); // 计算当前条目与上一条的时间差,生成中间缺失的条目 let nextTime = lastTime + step; while (nextTime < currentTime) { const formattedTime = new Date(nextTime).toISOString().slice(0, 23) + '000'; acc.push({ [timeHourKey]: formattedTime, [timeKey]: formattedTime, [countKey]: '0' }); nextTime += step; } // 加入当前条目 acc.push({...currentEntry}); } return acc; }, []); };
关键改动说明
- 明确字段名:直接指定三个字段的key,避免依赖API返回的字段顺序
- 正确的时间差计算:通过循环从最后一条的时间开始,逐步增加步长,直到追上当前条目的时间,生成所有缺失的中间条目
- 时间格式化:用
toISOString()确保生成的时间格式与原数据一致(保留三位毫秒) - 先处理初始状态:第一个条目直接加入结果集,后续条目先补全缺失项再加入自身
- 类型匹配:补全的
userCount设为字符串'0',与原数据类型保持一致
内容的提问来源于stack exchange,提问作者Igor Shmukler
相关产品推荐
相关产品推荐

