如何优化jq过滤器实现WebVTT转JSON?
优化jq过滤器实现WebVTT到JSON的转换
问题背景
需要将WebVTT文件转换为指定结构的JSON格式,现有jq过滤器能生成预期输出,但存在两个明显缺陷:
- 文件末尾必须添加额外换行才能输出最后一个cue
- 过滤器结构混乱,逐步修补导致可维护性差
示例WebVTT片段
WEBVTT dc9af9e8-8909-11ee-80b1-44850052397e 00:00:08.280 --> 00:00:08.880 Good morning Mr. Phelps. ddf6092c-8909-11ee-bc4f-44850052397e 00:00:09.720 --> 00:00:12.840 Your mission, should you choose to accept it, e03b2da2-8909-11ee-a212-44850052397e 00:17:52.052 --> 00:17:55.586 This tape will self-destruct in 5 seconds.
预期JSON输出
{ "filename": "<stdin>", "filetype": "WEBVTT", "cues": [ { "uid": "dc9af9e8-8909-11ee-80b1-44850052397e", "timeline": "00:00:08.280 --> 00:00:08.880", "subtitles": [ "Good morning Mr. Phelps." ] }, { "uid": "ddf6092c-8909-11ee-bc4f-44850052397e", "timeline": "00:00:09.720 --> 00:00:12.840", "subtitles": [ "Your mission, should you choose to", "accept it," ] }, { "uid": "e03b2da2-8909-11ee-a212-44850052397e", "timeline": "00:17:52.052 --> 00:17:55.586", "subtitles": [ "This tape will self-destruct in 5", "seconds." ] } ] }
优化后的jq过滤器
使用--null-input --raw-input参数执行以下过滤器:
def rtrim: rtrimstr("\r"); # 封装单个cue块的解析逻辑:UID行 → 时间线行 → 多行字幕 def parse_cue: . as $uid_line | (input | rtrim) as $timeline | [inputs | rtrim] | until(.[0] | test("^\\s*$") or . == []; . + [input | rtrim]) | map(select(length > 0)) | { uid: $uid_line | rtrim, timeline: $timeline, subtitles: . }; # 主处理逻辑 input | rtrim as $filetype | { filename: input_filename, filetype: $filetype, cues: [ foreach inputs as $line ( []; if $line | test("^[a-f0-9]{8}-") then . + [parse_cue] else . end; # 文件结束时强制输出所有已解析的cue,解决末尾换行依赖 if inputs == [] then . else . end ) ] }
优化说明
- 解决末尾换行依赖:通过在
foreach中检查inputs == [],在文件结束时自动输出所有已解析的cue,无需额外添加换行符。 - 模块化结构:将单个cue的解析逻辑封装到
parse_cue函数中,主逻辑更简洁,便于后续修改和扩展。 - 鲁棒性提升:使用
until循环读取字幕行,直到遇到空行或文件结束,自动过滤空行,避免无效内容混入结果。 - 保留跨平台兼容:保留
rtrim函数处理DOS格式的\r\n换行符,确保在不同系统下的兼容性。
内容的提问来源于stack exchange,提问作者rickhg12hs
相关产品推荐
相关产品推荐

