如何在循环中过滤不同withdrawal分组下的重复deposit ID?
过滤多组Withdrawal下重复Deposit ID的问题
我有一个包含多个withdrawal分组的列表,每个分组下对应一组deposit数据。需要实现:同一deposit ID只能出现在第一个包含它的withdrawal分组中,后续分组要过滤掉已出现过的deposit。withdrawal分组数量不固定,可能多于2个。尝试过基于lambda的过滤方法,但没得到预期结果。
示例输入
exampleList = [ { "withdrawal": { "amount": 250, "id": 70916631583, "date": "31-05-22 - 16:14:08", "paytype": "withdrawal" }, "deposit": [ { "id": 71018974368, "amount": 120, "date": "01-06-22 - 14:27:26", "paytype": "deposit" }, { "id": 71018971332, "amount": 100, "date": "01-06-22 - 14:27:23", "paytype": "deposit" } ] }, { "withdrawal": { "amount": 220, "id": 71019072820, "date": "01-06-22 - 14:28:40", "paytype": "withdrawal" }, "deposit": [ { "id": 71033338591, "amount": 100, "date": "01-06-22 - 17:03:19", "paytype": "deposit" }, { "id": 71033144597, "amount": 250, "date": "01-06-22 - 17:01:20", "paytype": "deposit" }, { "id": 71018974368, "amount": 120, "date": "01-06-22 - 14:27:26", "paytype": "deposit" }, { "id": 71018971332, "amount": 100, "date": "01-06-22 - 14:27:23", "paytype": "deposit" } ] } ]
预期输出
exampleOutputList = [ { "withdrawal": { "amount": 250, "id": 70916631583, "date": "31-05-22 - 16:14:08", "paytype": "withdrawal" }, "deposit": [ { "id": 71018974368, "amount": 120, "date": "01-06-22 - 14:27:26", "paytype": "deposit" }, { "id": 71018971332, "amount": 100, "date": "01-06-22 - 14:27:23", "paytype": "deposit" } ] }, { "withdrawal": { "amount": 220, "id": 71019072820, "date": "01-06-22 - 14:28:40", "paytype": "withdrawal" }, "deposit": [ { "id": 71033338591, "amount": 100, "date": "01-06-22 - 17:03:19", "paytype": "deposit" }, { "id": 71033144597, "amount": 250, "date": "01-06-22 - 17:01:20", "paytype": "deposit" } ] } ]
尝试过的代码及问题
之前尝试用循环记录已出现的ID,但逻辑错误,导致输出仍有重复ID:
listLen = len(exampleList) testList = [] if(listLen > 0): while listLen > 0: listLen -= 1 deposits = exampleList[listLen]['deposit'] withDrawal = exampleList[listLen]['withdrawal'] idList = [x['id'] for x in deposits] filterFromList = list(filter(lambda x:x['id'] not in testList, deposits)) testList.append({"withdrawal" : withDrawal,"deposit" : filterFromList}) print(testList)
输出结果里重复的deposit ID依然存在,因为testList存储的是整个分组对象,而非已出现的deposit ID集合,过滤条件完全不生效。
解决方案
核心是用一个集合来记录已经出现过的deposit ID,然后正序遍历每个withdrawal分组,过滤掉当前分组中已存在于集合的deposit,同时将新的deposit ID加入集合。
seen_deposit_ids = set() result = [] for group in exampleList: # 过滤当前分组中未出现过的deposit filtered_deposits = [d for d in group['deposit'] if d['id'] not in seen_deposit_ids] # 将当前分组中保留的deposit ID加入已见集合 seen_deposit_ids.update(d['id'] for d in filtered_deposits) # 构建结果分组 result.append({ "withdrawal": group['withdrawal'], "deposit": filtered_deposits }) # 打印格式化后的结果 import json print(json.dumps(result, indent=2))
代码说明
seen_deposit_ids是一个集合,用于快速判断ID是否已出现(集合的in操作时间复杂度为O(1),比列表高效)。- 正序遍历原列表,确保第一个出现的deposit ID被保留在对应的分组中。
- 过滤后,将当前分组保留的deposit ID全部加入集合,后续分组会自动过滤这些ID。
- 最后用
json.dumps格式化输出,方便查看结构。
运行这段代码后,就能得到符合预期的输出:每个deposit ID仅出现在第一个包含它的withdrawal分组中,后续分组不会重复出现。
内容的提问来源于stack exchange,提问作者Ozans
相关产品推荐
相关产品推荐

