You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scala中按规则对文件夹名字符串列表分组的实现需求

Scala 按文件夹名称正确分组(忽略末尾_1至_10后缀)

问题场景

现有一个包含文件夹名称的Scala List[String],部分名称末尾带有_1至_10的后缀,需要将同属一类的文件夹名分组到子列表中。初始列表如下:

val emp: List[String] = List(
  "customer_bal_history_1_36",
  "customer_bal_history_1_36_1",
  "customer_bal_history_1_36_2",
  "customer_bal_history_1_36_3",
  "customer_credit_history_37_72_1",
  "customer_credit_history_37_72_2",
  "customer_credit_history_37_72_3",
  "employee_1", 
  "employee_10", 
  "address",
  "pincode",
  "domain_1",
  "domain_2",
  "vehicle_1",
  "vehicle_2",
  "vendor_account_1",
  "vendor_account_2"
)

之前尝试用takeWhile(_ != '_')分组,会把customer_bal_history_1_36和customer_credit_history_37_72相关的文件夹错误分到同一组,需要将它们拆分为独立分组,得到符合预期的List[List[String]]。

解决方案:使用正则表达式分组

可以通过正则匹配去除末尾的_1至_10后缀,以此作为分组的key,实现精准分组。具体代码如下:

import scala.util.matching.Regex

// 定义正则:匹配末尾的_1到_10后缀,捕获前面的主体部分
val suffixPattern: Regex = """^(.*)_([1-9]|10)$""".r

// 提取分组key的函数
def getGroupKey(name: String): String = name match {
  case suffixPattern(base, _) => base
  case _ => name // 无后缀的名称直接作为key
}

// 执行分组并转换为目标列表格式
val grouped = emp.groupBy(getGroupKey).values.toList

代码说明

  1. 正则表达式:^(.*)_([1-9]|10)$ 精准匹配字符串末尾的_加1-9或10的后缀,(.*)捕获后缀之前的主体部分,作为分组的基准标识。
  2. 分组key提取:通过模式匹配,对带后缀的名称提取主体部分,无后缀的名称直接使用原字符串作为分组key。
  3. 分组执行:调用groupBy按提取的key完成分组,最后将分组后的Map值转换为List[List[String]]。

最终结果

执行上述代码后,得到的分组结果完全符合预期:

List(
  List(vehicle_1, vehicle_2),
  List(employee_1, employee_10),
  List(domain_1, domain_2),
  List(customer_bal_history_1_36, customer_bal_history_1_36_1, customer_bal_history_1_36_2, customer_bal_history_1_36_3),
  List(customer_credit_history_37_72_1, customer_credit_history_37_72_2, customer_credit_history_37_72_3),
  List(address),
  List(vendor_account_1, vendor_account_2),
  List(pincode)
)

内容的提问来源于stack exchange,提问作者gaurav mathur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 09:48:14