You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解析SRT文件时split()返回空数组,求正则及问题排查帮助

排查SRT解析中split()返回空数组的问题

我帮你拆解一下问题的核心:你的split()返回空数组,本质是你的SRT字符串格式不符合目标正则的匹配逻辑,同时先给你解释那个让你困惑的正则表达式到底在做什么。

先搞懂那个正则表达式

你提到的正则/(?:^|\n\n)\d+\n|\n+$/g是用来拆分标准SRT字幕的,拆解每个部分的作用:

  • (?:^|\n\n):这是个非捕获组,匹配字符串的开头,或者两个连续的换行符——这对应标准SRT里每个字幕块之间用空行分隔的规则。
  • \d+\n:匹配一个或多个数字(也就是字幕的序号),再加上一个换行符,用来定位每个字幕块的起始标记。
  • |:逻辑“或”,同时匹配后面的规则。
  • \n+$:匹配字符串末尾的一个或多个换行符,用来清理冗余的结尾空行。
  • g:全局匹配模式,会遍历整个字符串找所有符合的匹配项。

这个正则的设计前提是:你的SRT是标准格式——每个字幕块是序号\n时间轴\n文本\n\n的结构,而你的SRT字符串完全不符合这个要求,这才导致split后返回空数组。

你的SRT字符串的问题

看你定义的mySRT:

const mySRT = "0\ 00:00:01,414 --> 00:00:03,613\ All the single ladies\ 1\ 00:00:03,805 --> 00:00:05,904\ All the single ladies\ 2\ 00:00:06,118 --> 00:00:08,717\ All the single ladies\"

这里有两个致命问题:

  1. 你用\ (反斜杠加空格)来模拟换行,但这不是真正的换行符,JavaScript里真正的换行是\n,所以整个字符串是连续的一段文本,没有正则需要的\n或\n\n分隔符。
  2. 标准SRT里每个字幕块之间是空行(也就是\n\n),但你的字符串里完全没有这个结构。

解决方案步骤

1. 修复SRT字符串格式

把你的mySRT改成标准的SRT格式,用\n表示换行,空行分隔字幕块:

// 用模板字符串(反引号)可以直接写换行,更直观
const mySRT = `0
00:00:01,414 --> 00:00:03,613
All the single ladies

1
00:00:03,805 --> 00:00:05,904
All the single ladies

2
00:00:06,118 --> 00:00:08,717
All the single ladies`;

或者用\n拼接:

const mySRT = "0\n00:00:01,414 --> 00:00:03,613\nAll the single ladies\n\n1\n00:00:03,805 --> 00:00:05,904\nAll the single ladies\n\n2\n00:00:06,118 --> 00:00:08,717\nAll the single ladies";

2. 优化原解析函数(可选但推荐)

原函数的slice(1, -2)写法很脆弱,一旦SRT格式有小变动(比如末尾没有空行)就会出错,而且不支持多行文本的字幕。这里给你一个更健壮的版本:

function srtTimeToSeconds(time) {
  const match = time.match(/(\d\d):(\d\d):(\d\d),(\d\d\d)/);
  if (!match) throw new Error(`Invalid time format: ${time}`);
  const hours = +match[1], minutes = +match[2], seconds = +match[3], milliseconds = +match[4];
  return (hours * 60 * 60) + (minutes * 60) + seconds + (milliseconds / 1000);
}

function parseSrt(srt) {
  // 统一换行符为\n,清理多余的空白行,去掉首尾空白
  const normalizedSrt = srt.replace(/\r\n/g, '\n').replace(/\n{3,}/g, '\n\n').trim();
  // 按空行拆分每个字幕块
  const subtitleBlocks = normalizedSrt.split(/\n\n/);
  
  return subtitleBlocks.map(block => {
    // 拆分块内的行:序号行、时间轴行、剩下的是文本行(支持多行文本)
    const [indexLine, timeLine, ...textLines] = block.split(/\n/);
    // 解析时间轴
    const timeMatch = timeLine.match(/(\d\d:\d\d:\d\d,\d\d\d) --> (\d\d:\d\d:\d\d,\d\d\d)/);
    if (!timeMatch) throw new Error(`Invalid time line in block: ${block}`);
    
    return {
      index: parseInt(indexLine, 10),
      start: srtTimeToSeconds(timeMatch[1]),
      end: srtTimeToSeconds(timeMatch[2]),
      text: textLines.join('\n').trim()
    };
  });
}

3. 验证效果

现在调用函数就能得到正确的结果了:

const subtitles = parseSrt(mySRT);
console.log(subtitles);
// 输出:
// [
//   { index: 0, start: 1.414, end: 3.613, text: 'All the single ladies' },
//   { index: 1, start: 3.805, end: 5.904, text: 'All the single ladies' },
//   { index: 2, start: 6.118, end: 8.717, text: 'All the single ladies' }
// ]

内容的提问来源于stack exchange,提问作者LaTouv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:03:50