Scala中按规则对文件夹名字符串列表分组的实现需求
Scala 按文件夹名称正确分组(忽略末尾_1至_10后缀)
问题场景
现有一个包含文件夹名称的Scala List[String],部分名称末尾带有_1至_10的后缀,需要将同属一类的文件夹名分组到子列表中。初始列表如下:
val emp: List[String] = List( "customer_bal_history_1_36", "customer_bal_history_1_36_1", "customer_bal_history_1_36_2", "customer_bal_history_1_36_3", "customer_credit_history_37_72_1", "customer_credit_history_37_72_2", "customer_credit_history_37_72_3", "employee_1", "employee_10", "address", "pincode", "domain_1", "domain_2", "vehicle_1", "vehicle_2", "vendor_account_1", "vendor_account_2" )
之前尝试用takeWhile(_ != '_')分组,会把customer_bal_history_1_36和customer_credit_history_37_72相关的文件夹错误分到同一组,需要将它们拆分为独立分组,得到符合预期的List[List[String]]。
解决方案:使用正则表达式分组
可以通过正则匹配去除末尾的_1至_10后缀,以此作为分组的key,实现精准分组。具体代码如下:
import scala.util.matching.Regex // 定义正则:匹配末尾的_1到_10后缀,捕获前面的主体部分 val suffixPattern: Regex = """^(.*)_([1-9]|10)$""".r // 提取分组key的函数 def getGroupKey(name: String): String = name match { case suffixPattern(base, _) => base case _ => name // 无后缀的名称直接作为key } // 执行分组并转换为目标列表格式 val grouped = emp.groupBy(getGroupKey).values.toList
代码说明
- 正则表达式:
^(.*)_([1-9]|10)$精准匹配字符串末尾的_加1-9或10的后缀,(.*)捕获后缀之前的主体部分,作为分组的基准标识。 - 分组key提取:通过模式匹配,对带后缀的名称提取主体部分,无后缀的名称直接使用原字符串作为分组key。
- 分组执行:调用
groupBy按提取的key完成分组,最后将分组后的Map值转换为List[List[String]]。
最终结果
执行上述代码后,得到的分组结果完全符合预期:
List( List(vehicle_1, vehicle_2), List(employee_1, employee_10), List(domain_1, domain_2), List(customer_bal_history_1_36, customer_bal_history_1_36_1, customer_bal_history_1_36_2, customer_bal_history_1_36_3), List(customer_credit_history_37_72_1, customer_credit_history_37_72_2, customer_credit_history_37_72_3), List(address), List(vendor_account_1, vendor_account_2), List(pincode) )
内容的提问来源于stack exchange,提问作者gaurav mathur
相关产品推荐
相关产品推荐

