You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java/Kotlin:按类ID求含数据类的多个HashSet交集问题

我来帮你解决这个按id求Protein集合交集的问题——你遇到的核心问题应该是HashSet默认的相等性判断逻辑不符合你的需求,咱们一步步来解决:

问题根源分析

你定义的Protein类默认使用对象引用来判断相等性,而不是按id字段。HashSet的retainAll()(求交集)等操作完全依赖equals()和hashCode()方法,所以即使两个Protein实例的id相同,只要是不同对象,HashSet就会把它们当成不同元素,导致交集计算失败。

第一步:修正Protein类的相等性判断

我们需要重写equals()和hashCode()方法,让它们只基于id字段来判断相等性。这样HashSet就能正确识别出“同一个id的Protein是同一个元素”:

class Protein(val id: String, val score: Double, val molw: Double, val spc: Int) {
    override fun equals(other: Any?): Boolean {
        if (this === other) return true // 引用相同直接返回true
        if (other == null || javaClass != other.javaClass) return false // 类型不同返回false
        other as Protein
        return id == other.id // 仅比较id字段
    }

    override fun hashCode(): Int {
        return id.hashCode() // 哈希值仅基于id
    }
}
第二步:计算所有集合的交集

现在有两种高效的方式来计算6个HashSet的交集,你可以根据需求选择:

方式一:逐步缩小交集(简单直接)

以第一个集合为基准,依次和后续每个集合求交集,逐步筛选出所有集合共有的Protein:

fun findCommonProteins(proteinSets: ArrayList<HashSet<Protein>>): Set<Protein> {
    if (proteinSets.isEmpty()) return emptySet()
    
    // 初始化交集为第一个集合的副本(避免修改原集合)
    val commonProteins = HashSet(proteinSets[0])
    
    // 遍历剩余集合,不断求交集
    for (set in proteinSets.subList(1, proteinSets.size)) {
        commonProteins.retainAll(set)
        // 如果中途交集为空,提前终止循环提升效率
        if (commonProteins.isEmpty()) break
    }
    
    return commonProteins
}

方式二:先统计id出现次数(适合处理属性不一致的情况)

如果不同集合中同一个id对应的Protein其他属性(比如score、molw)可能不同,你可以先找出所有在6个集合中都出现的id,再根据id获取对应的Protein实例:

// 第一步:找出所有在所有集合中都出现的id
fun findCommonProteinIds(proteinSets: ArrayList<HashSet<Protein>>): Set<String> {
    val idOccurrenceCount = mutableMapOf<String, Int>()
    val totalSetCount = proteinSets.size

    for (set in proteinSets) {
        // 遍历每个集合中的Protein,统计id出现的集合次数
        set.forEach { protein ->
            idOccurrenceCount[protein.id] = idOccurrenceCount.getOrDefault(protein.id, 0) + 1
        }
    }

    // 筛选出在所有集合中都出现的id
    return idOccurrenceCount.filter { it.value == totalSetCount }.keys
}

// 第二步:根据id从第一个集合中获取对应的Protein实例(也可以从其他集合取,按需调整)
fun getCommonProteinsByIds(proteinSets: ArrayList<HashSet<Protein>>): Set<Protein> {
    val commonIds = findCommonProteinIds(proteinSets)
    return proteinSets[0].filter { it.id in commonIds }.toHashSet()
}
额外注意事项
  • 如果你的Protein是Kotlin的data class,默认会比较所有属性,所以即使是data class也需要手动重写equals()和hashCode()来只关注id。
  • 性能方面:两种方式都能高效处理数千元素的集合,方式二在交集较小时可能更快,因为不需要频繁修改集合;方式一更直观,适合属性完全一致的场景。

内容的提问来源于stack exchange,提问作者PeptideWitch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:33:31