如何将Pandas DataFrame重缩放至可配置的亚秒级均匀时间索引
问题描述
现有如下结构的Pandas DataFrame,存储音频转录短语的起止时间、内容和说话人标签:
>>> df start stop phrase speaker 0 1.22 3.13 im looking for shoes 1 1 5.13 5.45 okay 0 2 5.51 5.99 i can help with that 0 3 7.23 8.12 awesome 1 4 8.87 9.65 lets begin 0
可复现DataFrame的代码:
import pandas as pd data = [ (1.22, 3.13, 'im looking for shoes', 1), (5.13, 5.45, 'okay', 0), (5.51, 5.99, 'i can help with that', 0), (7.23, 8.12, 'awesome', 1), (8.87, 9.65, 'lets begin', 0) ] df = pd.DataFrame(data, columns=['start','stop','phrase','speaker'])
现有秒级对齐方案存在缺陷,同一秒内起止的短语会丢失数据,比如上述样例中的i can help with that就不会出现在输出结果中。需要实现支持灵活自定义0.1~0.25秒间隔的均匀时间索引对齐功能,且避免浮点数索引报错问题,输出格式参考如下(0.25秒间隔示例):
| frames | start | stop | phrase | speaker |
|---|---|---|---|---|
| 0.00 | NaN | NaN | NaN | NaN |
| 0.25 | NaN | NaN | NaN | NaN |
| ... | ... | ... | ... | ... |
| 1.25 | 1.22 | 3.13 | im looking for shoes | 1.0 |
| ... | ... | ... | ... | ... |
| 5.25 | 5.13 | 5.45 | okay | 0.0 |
| 5.50 | 5.51 | 5.99 | i can help with that | 0.0 |
| ... | ... | ... | ... | ... |
解决方案
如下实现支持灵活调整时间间隔,不会丢失同秒内的多短语数据,也不存在浮点数索引报错问题:
import pandas as pd import numpy as np def convert(df, step=0.25): # 计算时间范围上限,取最大结束时间向上取整 max_time = np.ceil(df['stop'].max()) # 生成全量均匀时间帧,保留两位小数避免浮点精度误差 frames = np.arange(0, max_time + step, step).round(2) full_df = pd.DataFrame({'frames': frames}) # 为原转录数据创建时间区间索引 df['interval'] = pd.IntervalIndex.from_arrays(df['start'], df['stop'], closed='both') # 匹配每个时间帧对应的转录条目 full_df['match_idx'] = full_df['frames'].apply( lambda x: df.index[df['interval'].apply(lambda interval: x in interval)] ).tolist() # 展开匹配结果(同一时间不会有多条重叠转录,每个帧最多匹配一条) full_df = full_df.explode('match_idx') # 关联原数据字段,清理冗余列 res = full_df.merge( df, left_on='match_idx', right_index=True, how='left' ).drop(columns=['match_idx', 'interval']) return res
使用示例
# 0.25秒间隔对齐 output_250ms = convert(df, step=0.25) # 0.1秒间隔对齐 output_100ms = convert(df, step=0.1)
逻辑说明
- 通过
step参数可自由调整时间间隔,支持0.1~0.25秒区间的所有测试需求 - 提前生成全量均匀时间帧,避免手动扩展行导致的丢数问题
- 用
IntervalIndex做区间匹配,逻辑简洁,不会漏掉同一秒内的多个短语 - 对时间帧做固定小数位舍入,从根源避免浮点数精度导致的匹配错误和索引报错
内容的提问来源于stack exchange,提问作者connor449
相关产品推荐
相关产品推荐

