带标签音频文件数据整理:按3秒块标记鸟类鸣声存在性
解决方案
可以通过以下步骤实现需求:
- 将原数据中的时间列转换为数值类型,方便后续计算;
- 为每个音频文件生成对应的3秒时间块序列;
- 检查每个时间块与原鸣声事件是否存在时间重叠,标记是否有鸣声。
完整实现代码
# 加载必要工具包 library(dplyr) library(tidyr) # 构建示例数据(与题目一致) id <- c("soundfile_1","soundfile_2","soundfile_3") sound_df<-data.frame(rep(id, each = 2), c("0","8.0","3.3","11.7","4.6","13.1"), c("3.2","14.1","3.8","12.8","5.9","14.8")) names(sound_df)[1] <- "soundfile" names(sound_df)[2] <- "sound_start" names(sound_df)[3] <- "sound_end" # 转换时间列为数值类型 sound_df <- sound_df %>% mutate(across(c(sound_start, sound_end), as.numeric)) # 定义音频总时长和块长度(真实数据将total_duration改为300即可) total_duration <- 15 block_length <- 3 # 生成每个音频的时间块序列 blocks_df <- sound_df %>% distinct(soundfile) %>% mutate( start = list(seq(0, total_duration - block_length, by = block_length)), end = list(seq(block_length, total_duration, by = block_length)) ) %>% unnest(c(start, end)) # 检查时间块与鸣声事件的重叠情况,生成最终结果 result_df <- blocks_df %>% left_join(sound_df, by = "soundfile") %>% mutate( # 判断时间块和鸣声事件是否有重叠 overlap = (start < sound_end) & (end > sound_start) ) %>% group_by(soundfile, start, end) %>% summarize( present = if_else(any(overlap, na.rm = TRUE), "yes", "no"), .groups = "drop" ) # 查看结果 print(result_df)
代码说明
- 类型转换:把
sound_start和sound_end从字符型转为数值型,避免时间计算出错; - 生成时间块:通过
seq()生成等间隔的3秒时间块,用unnest()将列表格式的时间序列展开为行; - 重叠判断:利用
(start < sound_end) & (end > sound_start)判断时间块与鸣声事件是否有交集,只要块内存在任意一次重叠,就标记为yes; - 分组汇总:按音频文件和时间块分组,汇总得到每个块的鸣声存在标记。
示例原数据
soundfile sound_start sound_end 1 soundfile_1 0 3.2 2 soundfile_1 8.0 14.1 3 soundfile_2 3.3 3.8 4 soundfile_2 11.7 12.8 5 soundfile_3 4.6 5.9 6 soundfile_3 13.1 14.8
预期输出结果
soundfile start end present 1 soundfile_1 0 3 yes 2 soundfile_1 3 6 yes 3 soundfile_1 6 9 yes 4 soundfile_1 9 12 yes 5 soundfile_1 12 15 yes 6 soundfile_2 0 3 no 7 soundfile_2 3 6 yes 8 soundfile_2 6 9 no 9 soundfile_2 9 12 yes 10 soundfile_2 12 15 yes 11 soundfile_3 0 3 no 12 soundfile_3 3 6 yes 13 soundfile_3 6 9 no 14 soundfile_3 9 12 no 15 soundfile_3 12 15 yes
内容的提问来源于stack exchange,提问作者davidj444
相关产品推荐
相关产品推荐

