如何在Scala中按列表子项分组统计对应评分总和
实现方案
核心思路
你之前直接groupBy(_._1)是把整条记录的完整分类列表当做分组键,不符合按单个分类拆分的需求。正确的处理逻辑是先把每条记录按分类拆分为「单个分类-对应评分」的元组,再按分类分组聚合即可。
针对示例数据的实现代码
val a = List(List(List("romance", "thriller"), 4), List(List("adventure", "thriller"), 3)) val result = a // 先格式统一化,把每条记录转成 (分类列表, 评分) 的强类型元组,避免List的Any类型问题 .map(record => (record(0).asInstanceOf[List[String]], record(1).asInstanceOf[Int])) // 按分类展开每条记录,得到所有(分类, 评分)的一维列表 .flatMap{case (categories, rating) => categories.map(cate => (cate, rating))} // 按分类分组 .groupBy(_._1) // 提取每个分组下的所有评分 .mapValues(_.map(_._2))
执行后result的输出就是你需要的结构:
Map( romance -> List(4), thriller -> List(4, 3), adventure -> List(3) )
适配Option类型的正式业务代码
你实际业务数据格式是(Option[List[String]], Option[Int]),需要补充空值过滤逻辑:
// 假设你的实际业务数据格式如下 val businessData: List[(Option[List[String]], Option[Int])] = List( (Some(List("romance", "thriller")), Some(4)), (Some(List("adventure", "thriller")), Some(3)), (None, Some(2)), // 分类为空的无效数据会被过滤 (Some(List("science")), None) // 评分为空的无效数据会被过滤 ) val formalResult = businessData // 过滤掉分类为空 或 评分为空的无效记录 .collect{case (Some(cates), Some(rating)) => (cates, rating)} // 按分类拆分记录 .flatMap{case (cates, rating) => cates.map(cate => (cate, rating))} // 分组聚合 .groupBy(_._1) .mapValues(_.map(_._2))
结果调整说明
如果你需要最终评分是Tuple格式而非List,可以在mapValues阶段补充转换:
- 单元素转单值Tuple:
mapValues(list => if(list.size == 1) Tuple1(list.head) else list) - 统一转Tuple:
mapValues(_.toTuple)(Scala 2.13+支持)
内容的提问来源于stack exchange,提问作者NanoBit
相关产品推荐
相关产品推荐

