按空行分割SRT字符串为指定数量对话块的Kotlin实现方案
问题
我有SRT格式的字符串,需要分割成多个块,每个块包含指定数量的SRT对话条目。目前使用按行数分割的Kotlin函数,但出现内容截断、块数计算错误的问题(比如6664行的SRT被分割成28块,逻辑错误),希望改为按空行分割,确保每个块包含指定数量(如400个)的完整SRT条目。
现有分割函数:
fun divideStringIntoChunks( input: String, chunkSize: Int, fromLocal: String, intoLocale: String ): ArrayList<String> { val formattedInput = formatString(input) val lines = formattedInput.split("%0A") val chunks = lines.chunked(chunkSize) val result = ArrayList<String>() for (chunk in chunks) { val joinChunk = chunk.joinToString("%0A") val emptyLineAfterFormatted = joinChunk + "\n" result.add("https://translate.google.com/?sl=$fromLocal&tl=$intoLocale&text=$emptyLineAfterFormatted&op=translate") } return result }
编码函数:
fun formatString(input: String): String { val encodedString = URLEncoder.encode(input, "UTF-8") return encodedString.replace("+", "%20").replace("%2C", ",") }
SRT示例:
1474 02:13:11,225 --> 02:13:13,425 Thank you. I love when you call me Princess. 1475 02:13:14,762 --> 02:13:16,130 You work across the street? 1476 02:13:16,797 --> 02:13:17,799 Yeah. 1477 02:13:18,464 --> 02:13:20,098 Diamond exchange?
解决方案
核心思路是先将SRT字符串按空行分割成独立的对话条目,再将这些条目按指定数量分块,最后处理编码并生成翻译链接。修改后的代码如下:
fun divideSrtIntoEntryChunks( input: String, entriesPerChunk: Int, fromLocale: String, intoLocale: String ): List<String> { // 按空行分割成单个SRT条目,兼容\n\n和\r\n\r\n两种换行格式 val srtEntries = input.split(Regex("(?<=\n)\n|(?<=\r\n)\r\n")) .filter { it.isNotBlank() } // 过滤首尾或中间的空条目 // 将完整条目按指定数量分块 val entryChunks = srtEntries.chunked(entriesPerChunk) // 处理每个块,生成对应的翻译链接 return entryChunks.map { chunk -> // 用空行连接块内条目,还原SRT格式 val chunkContent = chunk.joinToString("\n\n") // 编码内容以适配URL查询参数 val encodedContent = formatString(chunkContent) // 构造翻译链接 "https://translate.google.com/?sl=$fromLocale&tl=$intoLocale&text=$encodedContent&op=translate" } } // 保留原有的编码函数 fun formatString(input: String): String { val encodedString = URLEncoder.encode(input, "UTF-8") return encodedString.replace("+", "%20").replace("%2C", ",") }
关键说明
- 精准分割条目:使用正则
(?<=\n)\n|(?<=\r\n)\r\n匹配SRT中分隔条目的空行,确保每个分割结果都是完整的「编号+时间轴+字幕内容」组合,彻底避免单个条目被截断的问题。 - 清理无效条目:通过
filter { it.isNotBlank() }过滤掉首尾空行或多余空行产生的无效条目,保证条目列表的整洁性。 - 按条目分块:直接对完整条目列表调用
chunked,确保每个块严格包含指定数量的SRT对话,解决原按行数分割导致的块数计算错误问题。 - 格式还原:每个块内的条目用空行连接,还原标准SRT格式后再编码,保证翻译时的格式正确性。
内容的提问来源于stack exchange,提问作者Roony
相关产品推荐
相关产品推荐

