如何识别并分组符合条件的重复交易记录?求正确代码实现
重复交易记录识别问题
需求说明
找出所有满足以下条件的重复交易记录:
- sourceAccount、targetAccount、category、amount完全相同;
- 连续交易之间的时间差小于1分钟。
需将符合条件的交易按组划分:
- 组内交易按时间升序排列;
- 各组按组内首条交易的时间升序排列。
数据集
[ { "id": 3, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:34:30.000Z" }, { "id": 1, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:33:00.000Z" }, { "id": 6, "sourceAccount": "A", "targetAccount": "C", "amount": 250, "category": "other", "time": "2018-03-02T10:33:05.000Z" }, { "id": 4, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:36:00.000Z" }, { "id": 2, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:33:50.000Z" }, { "id": 5, "sourceAccount": "A", "targetAccount": "C", "amount": 250, "category": "other", "time": "2018-03-02T10:33:00.000Z" } ]
预期输出
[ [ { "id": 1, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:33:00.000Z" }, { "id": 2, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:33:50.000Z" }, { "id": 3, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:34:30.000Z" } ], [ { "id": 5, "sourceAccount": "A", "targetAccount": "C", "amount": 250, "category": "other", "time": "2018-03-02T10:33:00.000Z" }, { "id": 6, "sourceAccount": "A", "targetAccount": "C", "amount": 250, "category": "other", "time": "2018-03-02T10:33:05.000Z" } ] ]
尝试的代码
let transactions = [] var clean = records.filter((arr, index, self) => index === self.findIndex((t) => (t.sourceAccount === arr.sourceAccount && t.targetAccount === arr.targetAccount && t.category === arr.category && t.amount === arr.amount))) if(clean.length > 1){ // array contains duplicate elements. transactions.push(clean) } return transactions
实际输出
[ [ { "id": 3, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:34:30.000Z" }, { "id": 6, "sourceAccount": "A", "targetAccount": "C", "amount": 250, "category": "other", "time": "2018-03-02T10:33:05.000Z" }, { "id": 4, "sourceAccount": "Aw", "targetAccount": "D", "amount": 1002, "category": "eating_out2", "time": "2018-03-02T10:36:00.000Z" } ] ]
问题分析
原代码存在以下核心问题:
- 使用
filter+findIndex的逻辑会保留每个特征组的第一个元素,反而过滤掉同组其他交易,与需求完全相悖; - 未对交易按时间排序,无法正确判断连续交易的时间差;
- 完全忽略了「连续交易时间差小于1分钟」的核心条件;
- 分组逻辑错误,将不同特征的交易混为一组。
正确代码实现
function findDuplicateTransactions(records) { // 1. 先将所有交易按时间升序排序 const sortedRecords = [...records].sort((a, b) => new Date(a.time) - new Date(b.time)); // 2. 按sourceAccount、targetAccount、category、amount分组 const featureGroups = {}; sortedRecords.forEach(record => { // 生成唯一分组键 const groupKey = `${record.sourceAccount}-${record.targetAccount}-${record.category}-${record.amount}`; if (!featureGroups[groupKey]) { featureGroups[groupKey] = []; } featureGroups[groupKey].push(record); }); // 3. 拆分出连续时间差小于1分钟的有效子组 const validGroups = []; Object.values(featureGroups).forEach(group => { if (group.length < 2) return; // 跳过单条交易的组 let currentSubGroup = [group[0]]; for (let i = 1; i < group.length; i++) { const prevTimestamp = new Date(currentSubGroup.at(-1).time).getTime(); const currTimestamp = new Date(group[i].time).getTime(); const diffMinutes = (currTimestamp - prevTimestamp) / (1000 * 60); if (diffMinutes < 1) { currentSubGroup.push(group[i]); } else { // 时间差超标,保存当前有效子组并重置 if (currentSubGroup.length >= 2) { validGroups.push(currentSubGroup); } currentSubGroup = [group[i]]; } } // 处理最后一个子组 if (currentSubGroup.length >= 2) { validGroups.push(currentSubGroup); } }); // 4. 按组内首条交易时间升序排列结果 return validGroups.sort((a, b) => new Date(a[0].time) - new Date(b[0].time)); } // 测试调用 const records = [ { "id": 3, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:34:30.000Z" }, { "id": 1, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:33:00.000Z" }, { "id": 6, "sourceAccount": "A", "targetAccount": "C", "amount": 250, "category": "other", "time": "2018-03-02T10:33:05.000Z" }, { "id": 4, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:36:00.000Z" }, { "id": 2, "sourceAccount": "A", "targetAccount": "B", "amount": 100, "category": "eating_out", "time": "2018-03-02T10:33:50.000Z" }, { "id": 5, "sourceAccount": "A", "targetAccount": "C", "amount": 250, "category": "other", "time": "2018-03-02T10:33:00.000Z" } ]; console.log(findDuplicateTransactions(records));
代码说明
- 排序:先对交易按时间升序排列,确保后续能正确判断连续交易的时间差;
- 特征分组:用拼接字符串生成唯一键,将特征完全一致的交易归为一组;
- 拆分有效子组:遍历每个特征组,逐个判断交易间的时间差,小于1分钟则归入当前子组,否则结束当前子组并重置;
- 结果整理:过滤掉单条交易的无效组,最后按子组首条交易时间排序,得到符合要求的结果。
内容的提问来源于stack exchange,提问作者Learn with Panda
相关产品推荐
相关产品推荐

