You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

VTT字幕内嵌时间戳与文本提取:正则匹配实现方案咨询

解决VTT内嵌时间戳与文本提取的PHP实现方案

我明白你需要从复杂的VTT字幕里提取内嵌时间戳和对应文本的需求,你之前写的正则<([^;]*)>确实太宽泛了,会匹配到所有尖括号里的内容(包括颜色标签),咱们来调整一下,用更精准的正则和PHP代码来实现你的目标。

完整实现代码

<?php
// 示例VTT内容(你可以替换成读取文件的内容,比如用file_get_contents读取本地VTT文件)
$vttContent = '00:01:03.500 --> 00:01:03.510 align:start position:0% <c.colorCCCCCC>fourth guess it came from a broken</c><c.colorE5E5E5> home </c> 00:01:03.510 --> 00:01:08.140 align:start position:0% <c.colorCCCCCC>fourth guess it came from a broken</c><c.colorE5E5E5> home a<00:01:04.580><c> father</c><00:01:05.580><c> not</c><00:01:05.820><c> being</c><00:01:05.880><c> there</c><00:01:06.890><c> my</c><00:01:07.890><c> mother</c></c> 00:01:08.140 --> 00:01:08.150 align:start position:0% a<c.colorE5E5E5> father not being there my mother </c> 00:01:08.150 --> 00:01:13.429 align:start position:0% a<c.colorE5E5E5> father not being there my mother</c> <c.colorE5E5E5>getting<00:01:09.150><c> married</c><00:01:09.630><c> and</c><00:01:11.360><c> the</c><00:01:12.360><c> abuse</c></c><c.colorCCCCCC><00:01:12.659><c> started</c><00:01:13.049><c> at</c></c>';

// 初始化结果数组
$timestamps = [];
$texts = [];

// 第一步:匹配所有内嵌的 <时间戳><c>文本</c> 结构
$inlinePattern = '/<(\d{2}:\d{2}:\d{2}\.\d{3})><c>(.*?)<\/c>/';
preg_match_all($inlinePattern, $vttContent, $inlineMatches, PREG_SET_ORDER);

// 收集这些内嵌的时间戳和文本
foreach ($inlineMatches as $match) {
    $timestamps[] = $match[1];
    $texts[] = trim($match[2]);
}

// 第二步:处理开头的第一个时间戳对应的完整文本(比如00:01:03.510对应的"fourth guess it came from a broken home")
$firstInlineTsPos = strpos($vttContent, '<00:');
if ($firstInlineTsPos !== false) {
    // 截取第一个内嵌时间戳前的内容
    $preTextSegment = substr($vttContent, 0, $firstInlineTsPos);
    
    // 清理颜色标签和VTT行头的时间范围、属性
    $cleanPreText = preg_replace('/<\/?c\.[a-fA-F0-9]{6}>/', '', $preTextSegment);
    $cleanPreText = preg_replace('/^\d{2}:\d{2}:\d{2}\.\d{3} --> \d{2}:\d{2}:\d{2}\.\d{3} align:[^%]+% /', '', $cleanPreText);
    $cleanPreText = trim($cleanPreText);
    
    // 提取开头的目标时间戳(00:01:03.510)
    preg_match('/(\d{2}:\d{2}:\d{2}\.\d{3}) -->/', $vttContent, $firstTsMatch);
    if (!empty($firstTsMatch[1])) {
        // 将开头的时间戳和文本插入数组最前面,保证顺序正确
        array_unshift($timestamps, $firstTsMatch[1]);
        array_unshift($texts, $cleanPreText);
    }
}

// 输出结果(你可以根据需求调整输出方式,比如返回数组或写入文件)
echo "时间戳数组:\n";
print_r($timestamps);
echo "\n文本数组:\n";
print_r($texts);
?>

代码说明

  1. 正则精准匹配:

    • 用/<(\d{2}:\d{2}:\d{2}\.\d{3})><c>(.*?)<\/c>/专门匹配内嵌的时间戳+文本结构,避开颜色标签的干扰。
    • 对于开头的时间戳,通过截取+清理的方式,提取出对应的完整字幕文本,同时去掉VTT行的冗余属性。
  2. 结果顺序保证:

    • 先收集所有中间的内嵌时间戳和文本,再把开头的时间戳和文本插入数组头部,完全符合你需要的顺序。
    • 用trim()清理文本中的多余空格,让结果更整洁。

运行结果

执行代码后,你会得到和需求完全一致的两个数组:

  • 时间戳数组包含所有目标时间点,顺序正确
  • 文本数组对应每个时间点的字幕内容

内容的提问来源于stack exchange,提问作者Peter Friedlander

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:54:58