You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Gremlin查询找出重复记录、更新isDuplicate属性并过滤单条记录

Fixing Duplicate Vertex Detection & Attribute Update in Gremlin

Let's work through this step by step to solve your problem—we need to first pinpoint vertices with duplicate customerId values, then mark all of them by setting their isDuplicate property to "yes". Your original query pulls all customerId groups but includes single-entry ones, so we'll adjust it to target only duplicates, then apply the update.

Step 1: Isolate Duplicate customerId Values

Instead of using group() to collect vertices per ID, we'll use groupCount() to tally how many times each customerId appears. Then we filter out IDs that only show up once, leaving us with just the duplicates:

g.V()
  .has("customerId")
  .groupCount().by("customerId")
  .unfold()
  .filter{ it.value > 1 }
  .map{ it.key }
  .toList()

This returns a clean list of customerId values that exist on 2 or more vertices.

Step 2: Update All Duplicate Vertices

Now we can use that list of duplicate IDs to update every matching vertex. You have two clean options here:

Option 1: Use a Variable (More Readable)

First capture the duplicate IDs, then reference them to update vertices:

// Save duplicate customer IDs to a variable
def duplicateCustomerIds = g.V()
  .has("customerId")
  .groupCount().by("customerId")
  .unfold()
  .filter{ it.value > 1 }
  .map{ it.key }
  .toList()

// Update all vertices with those duplicate IDs
g.V()
  .has("customerId", within(duplicateCustomerIds))
  .property("isDuplicate", "yes")
  .iterate()

Option 2: Single-Chained Query

If you prefer to handle it in one continuous query without a variable:

g.V()
  .has("customerId")
  .groupCount().by("customerId")
  .unfold()
  .filter{ it.value > 1 }
  .map{ it.key }
  .flatMap{ g.V().has("customerId", it) }
  .property("isDuplicate", "yes")
  .iterate()

Quick Tips:

  • Use iterate() instead of toList() for update operations—it tells Gremlin to execute the write action without returning unnecessary results, which is more efficient.
  • The within() step ensures we match every vertex with a duplicate customerId, so no instances of the duplicate get missed.

内容的提问来源于stack exchange,提问作者Ksharapani Wickramarathne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:45:23