You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

寻求dplyr方案:在语音转录数据中插入Utterance间时间间隔行

使用dplyr实现语音转录数据的时间间隔行插入

需求说明

我正在处理语音转录数据,现有数据如下:

Utterance                       Starttime_ms Endtime_ms
  <chr>                                  <dbl>      <dbl>
1 on this                                  210        780
2 okay                                    3403       3728
3 cool thanks everyone um                 4221       5880
4 so yes in terms of our projects         5910      11960
5 let's have a look so the               11980      13740
6 LGBTQ plus                             13813      16110

需要在每个Utterance行之后插入新行,用来表示该行与前一个Utterance之间的时间间隔,期望输出如下:

Utterance                       Starttime_ms Endtime_ms
  <chr>                                  <dbl>      <dbl>
1 on this                                  210        780
  NA                                       780       3403
2 okay                                    3403       3728
  NA                                      3728       4221
3 cool thanks everyone um                 4221       5880
  NA                                      5880       5910
4 so yes in terms of our projects         5910      11960
  NA                                     11960      11980
5 let's have a look so the               11980      13740
  NA                                     13740      13813
6 LGBTQ plus                             13813      16110

我已经掌握了使用data.table实现该需求的方法:

library(data.table)
unq <- c(0, sort(unique(setDT(df)[, c(Starttime_ms, Endtime_ms)])))
df <- df[.(unq[-length(unq)], unq[-1]), on=c("Starttime_ms", "Endtime_ms")]

现寻求使用dplyr包实现该需求的解决方案,数据结构如下:

df <-   structure(list(Utterance = c("on this", "okay", "cool thanks everyone um", 
                                     "so yes in terms of our projects", 
                                     "let's have a look so the", "LGBTQ plus"), Starttime_ms = c(210, 
                                                                                                 3403, 4221, 5910, 11980, 13813), Endtime_ms = c(780, 3728, 5880, 
                                                                                                                                                 11960, 13740, 16110)), row.names = c(NA, -6L), class = c("tbl_df", 
                                                                                                                                                                                                          "tbl", "data.frame"))

dplyr解决方案

方法一:基于时间区间生成与左连接

library(dplyr)

# 提取所有时间点并生成连续的时间区间
time_intervals <- df %>%
  select(Starttime_ms, Endtime_ms) %>%
  tidyr::pivot_longer(cols = everything(), values_to = "time") %>%
  pull(time) %>%
  c(0, .) %>%
  sort() %>%
  {tibble(Starttime_ms = head(., -1), Endtime_ms = tail(., -1))}

# 左连接原始数据,自动填充NA行
result <- time_intervals %>%
  left_join(df, by = c("Starttime_ms", "Endtime_ms")) %>%
  select(Utterance, Starttime_ms, Endtime_ms)

print(result)

方法二:生成间隔行后合并排序

library(dplyr)

df %>%
  # 标记下一个 utterance 的起始时间
  mutate(next_start = lead(Starttime_ms)) %>%
  # 生成时间间隔对应的NA行
  bind_rows(
    tibble(
      Utterance = NA,
      Starttime_ms = .$Endtime_ms,
      Endtime_ms = .$next_start
    ) %>% filter(!is.na(Endtime_ms))
  ) %>%
  # 按起始时间排序
  arrange(Starttime_ms) %>%
  # 移除辅助列
  select(-next_start)

这段代码逻辑:先给每行标记下一个语音段的起始时间,再基于当前行的结束时间和下一行的起始时间生成空语音段行,最后合并所有行并按时间排序即可得到目标结果。

内容的提问来源于stack exchange,提问作者Chris Ruehlemann

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 18:53:18