You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ruby中从哈希数组提取基于identifier键的唯一元素并保留存储索引

我明白你要做的是从一组包含哈希和索引的子数组里,筛选出那些带有新的唯一identifier的项,同时保留它们完整的子数组结构。我来给你梳理下实现思路和代码:

实现思路

核心是跟踪已经出现过的identifier,避免重复收集仅包含已见过id的子数组:

  • 用Ruby的Set来存储已处理的identifier(比数组查找效率更高)
  • 遍历每个子数组时:
    1. 分离出子数组中的哈希对象(排除最后一个索引元素)
    2. 提取这些哈希的identifier值
    3. 检查是否存在未被记录过的identifier
    4. 如果有新id,就保留当前子数组,并把新id加入已记录集合

代码实现

require 'set'

# 你的输入数据
input = [
  [{ "identifier"=>"A", "inclusion"=>"true", "name"=>"FGFR2", "specific"=>"false" }, 0],
  [{ "identifier"=>"A", "inclusion"=>"true", "name"=>"FGFR2", "specific"=>"false" }, 1],
  [{ "identifier"=>"B", "inclusion"=>"true", "name"=>"ALK", "specific"=>"false" }, { "identifier"=>"A", "inclusion"=>"true", "name"=>"FGFR2", "specific"=>"false" }, 4],
  [{ "identifier"=>"A", "inclusion"=>"true", "name"=>"FGFR2", "specific"=>"false" }, 5]
]

seen_ids = Set.new
result = []

input.each do |sub_arr|
  # 提取子数组中除最后一个索引外的所有哈希对象
  hashes = sub_arr[0...-1]
  # 获取当前子数组的所有identifier
  current_ids = hashes.map { |h| h["identifier"] }
  # 找出未见过的新id
  new_ids = current_ids - seen_ids.to_a

  unless new_ids.empty?
    result << sub_arr
    # 将新id加入已见集合,避免后续重复收集
    seen_ids.merge(new_ids)
  end
end

# 打印结果
p result

运行结果

这段代码执行后,会输出你期望的结果:

[
  [{ "identifier"=>"A", "inclusion"=>"true", "name"=>"FGFR2", "specific"=>"false" }, 0],
  [{ "identifier"=>"B", "inclusion"=>"true", "name"=>"ALK", "specific"=>"false" }, { "identifier"=>"A", "inclusion"=>"true", "name"=>"FGFR2", "specific"=>"false" }, 4]
]

补充说明

  • 选择Set而不是数组来存储已见id,是因为它的查找和合并操作时间复杂度更低,数据量大时优势更明显
  • sub_arr[0...-1]的写法可以兼容子数组中有多个哈希的情况(比如第三个子数组有两个哈希),不用硬写索引位置
  • 只有当子数组包含至少一个新id时才会被保留,完美匹配你要的去重逻辑

内容的提问来源于stack exchange,提问作者Robin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 19:57:47