Scala中带前缀保留的字符串拆分及Case Class映射实现
Great question! The problem with your initial split approach is that it discards the Z01-Z0D identifiers we need to map to specific case classes. Let's solve this by using regex capture groups to preserve these delimiters, then process the split results into structured data.
Step 1: Adjust the Regex to Retain Delimiters
Instead of splitting on the |Z0X| pattern directly, wrap it in a capture group (()). This tells Scala's split method to include the delimiter itself in the resulting array, rather than discarding it.
优化后的正则表达式:
val splitRegex = """(\|Z0[1-9A-D]\|)"""
[1-9A-D]is a simplified alternative to your original list (matches digits 1-9 and letters A-D)- The capture group ensures the delimiter (e.g.,
|Z01|) is kept as a separate element in the split array.
Step 2: Process the Split Results
When we split the string with this regex, the array will look like:
Array( "X|blnk_1|...|blnk_8| ", // X section "|Z01|", // First Z01 delimiter "Str1|01|...|[00]|]|", // First Z01 content "|Z01|", // Second Z01 delimiter "Str2|02|...|[01]|]|", // Second Z01 content "|Z02|", // Z02 delimiter "02|z2|Str|Str|" // Z02 content )
We can clean this up to get your desired formatted output, then map each section to its case class:
val str = "X|blnk_1|blnk_2|blnk_3|blnk_4|time1|time2|blnk_5|blnk_6|blnk_7|blnk_8| |Z01|Str1|01|001|NE]|[HEX1|HEX2]|[NA|001:1000|123:456|[00]|]|Z01|Str2|02|002|NE]|[HEX3|HEX4]|[NA|002:1001|234:456|[01]|]|Z02|02|z2|Str|Str|" // Split with capture group to retain delimiters val splitParts = str.split("""(\|Z0[1-9A-D]\|)""") // Process the X section (first element) val xSection = splitParts.head.trim.replaceAll("\\|$", "") // Clean trailing space and | println(xSection) // Output: X|blnk_1|blnk_2|blnk_3|blnk_4|time1|time2|blnk_5|blnk_6|blnk_7|blnk_8 // Process Z0X sections: group delimiters with their content val zSections = splitParts.tail.grouped(2).map { case Array(delimiter, content) => val identifier = delimiter.replaceAll("\\|", "") // Convert |Z01| to Z01 s"$identifier|$content" }.toList // Print formatted Z sections (matches your desired output) zSections.foreach(println) // Output: // Z01|Str1|01|001|NE]|[HEX1|HEX2]|[NA|001:1000|123:456|[00]|]| // Z01|Str2|02|002|NE]|[HEX3|HEX4]|[NA|002:1001|234:456|[01]|]| // Z02|02|z2|Str|Str|
Step 3: Map to Case Classes
Now we can parse each section into its corresponding case class. First, define your case classes:
// Case class for the X section case class X( blnk1: String, blnk2: String, blnk3: String, blnk4: String, time1: String, time2: String, blnk5: String, blnk6: String, blnk7: String, blnk8: String ) // Case class for Z01 sections case class Z01( str: String, code1: String, code2: String, ne: String, hexPair: (String, String), details: (String, String, String, String) ) // Case class for Z02 sections case class Z02( code: String, z2: String, str1: String, str2: String, str3: String )
Then parse each section:
// Parse X instance val xFields = xSection.split("\\|").tail // Skip the leading "X" val xInstance = X( xFields(0), xFields(1), xFields(2), xFields(3), xFields(4), xFields(5), xFields(6), xFields(7), xFields(8), xFields(9) ) // Parse Z01 instances val z01Instances = zSections.filter(_.startsWith("Z01|")).map { section => val fields = section.split("\\|").tail // Skip leading "Z01" val ne = fields(3).replace("]", "") // Clean "NE]" to "NE" val hexParts = fields(4).stripPrefix("[").stripSuffix("]").split("\\|") val detailParts = fields(5).stripPrefix("[").stripSuffix("]").split("\\|") Z01( fields(0), fields(1), fields(2), ne, (hexParts(0), hexParts(1)), (detailParts(0), detailParts(1), detailParts(2), detailParts(3)) ) } // Parse Z02 instances val z02Instances = zSections.filter(_.startsWith("Z02|")).map { section => val fields = section.split("\\|").tail // Skip leading "Z02" Z02(fields(0), fields(1), fields(2), fields(3), fields(4)) }
Key Takeaways
- Use regex capture groups in
splitto retain delimiters instead of discarding them. - Group delimiter elements with their corresponding content to rebuild complete Z0X sections.
- Clean and parse each section into case classes by splitting fields and handling special formatting (like
[]brackets).
内容的提问来源于stack exchange,提问作者Robert Knox

