You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas DataFrame重缩放至可配置的亚秒级均匀时间索引

问题描述

现有如下结构的Pandas DataFrame,存储音频转录短语的起止时间、内容和说话人标签:

>>> df
   start  stop                phrase  speaker
0   1.22  3.13  im looking for shoes        1
1   5.13  5.45                  okay        0
2   5.51  5.99  i can help with that        0
3   7.23  8.12               awesome        1
4   8.87  9.65            lets begin        0

可复现DataFrame的代码:

import pandas as pd
data = [
    (1.22, 3.13, 'im looking for shoes', 1),
    (5.13, 5.45, 'okay', 0),
    (5.51, 5.99, 'i can help with that', 0),
    (7.23, 8.12, 'awesome', 1),
    (8.87, 9.65, 'lets begin', 0)
]
df = pd.DataFrame(data, columns=['start','stop','phrase','speaker'])

现有秒级对齐方案存在缺陷,同一秒内起止的短语会丢失数据,比如上述样例中的i can help with that就不会出现在输出结果中。需要实现支持灵活自定义0.1~0.25秒间隔的均匀时间索引对齐功能,且避免浮点数索引报错问题,输出格式参考如下(0.25秒间隔示例):

framesstartstopphrasespeaker
0.00NaNNaNNaNNaN
0.25NaNNaNNaNNaN
...............
1.251.223.13im looking for shoes1.0
...............
5.255.135.45okay0.0
5.505.515.99i can help with that0.0
...............
解决方案

如下实现支持灵活调整时间间隔,不会丢失同秒内的多短语数据,也不存在浮点数索引报错问题:

import pandas as pd
import numpy as np

def convert(df, step=0.25):
    # 计算时间范围上限,取最大结束时间向上取整
    max_time = np.ceil(df['stop'].max())
    # 生成全量均匀时间帧,保留两位小数避免浮点精度误差
    frames = np.arange(0, max_time + step, step).round(2)
    full_df = pd.DataFrame({'frames': frames})
    # 为原转录数据创建时间区间索引
    df['interval'] = pd.IntervalIndex.from_arrays(df['start'], df['stop'], closed='both')
    # 匹配每个时间帧对应的转录条目
    full_df['match_idx'] = full_df['frames'].apply(
        lambda x: df.index[df['interval'].apply(lambda interval: x in interval)]
    ).tolist()
    # 展开匹配结果(同一时间不会有多条重叠转录,每个帧最多匹配一条)
    full_df = full_df.explode('match_idx')
    # 关联原数据字段,清理冗余列
    res = full_df.merge(
        df, left_on='match_idx', right_index=True, how='left'
    ).drop(columns=['match_idx', 'interval'])
    return res

使用示例

# 0.25秒间隔对齐
output_250ms = convert(df, step=0.25)
# 0.1秒间隔对齐
output_100ms = convert(df, step=0.1)

逻辑说明

  • 通过step参数可自由调整时间间隔,支持0.1~0.25秒区间的所有测试需求
  • 提前生成全量均匀时间帧,避免手动扩展行导致的丢数问题
  • 用IntervalIndex做区间匹配,逻辑简洁,不会漏掉同一秒内的多个短语
  • 对时间帧做固定小数位舍入,从根源避免浮点数精度导致的匹配错误和索引报错

内容的提问来源于stack exchange,提问作者connor449

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 23:30:05