在Pandas DataFrame中插入新行,生成10分钟间隔时间序列并保留数据
解决方案:将小时间隔数据转换为10分钟间隔
这是一个很常见的时间序列扩展需求,用Pandas可以轻松实现,这里给你两种可行的方案:
方法一:手动生成时间序列并展开
这种方法直接为每一行原始数据生成包含6个10分钟间隔的时间点,然后展开成新行,适合需要明确控制时间点生成逻辑的场景。
步骤代码:
import pandas as pd # 1. 构造原始DataFrame(你可以直接读取自己的数据源,这里用你提供的示例数据) data = [ [1, "2016-10-01 01:00:00", 1014.7, 23.6], [2, "2016-10-01 02:00:00", 1014.3, 23.6], [3, "2016-10-01 03:00:00", 1014.3, 23.8], [4, "2016-10-01 04:00:00", 1014.3, 23.8], [5, "2016-10-01 05:00:00", 1014.4, 24.3], [6, "2016-10-01 06:00:00", 1014.9, 24.6], [7, "2016-10-01 07:00:00", 1015.6, 25.7], [8, "2016-10-01 08:00:00", 1015.8, 26], [9, "2016-10-01 09:00:00", 1016.3, 27.3], [10, "2016-10-01 10:00:00", 1016.5, 25.8], [11, "2016-10-01 11:00:00", 1016.6, 26], [12, "2016-10-01 12:00:00", 1016.6, 27.3] ] df = pd.DataFrame(data, columns=["id", "timestamp", "pressure", "temp"]) # 2. 将timestamp转换为datetime类型(必须步骤,否则无法生成时间序列) df["timestamp"] = pd.to_datetime(df["timestamp"]) # 3. 为每一行生成6个10分钟间隔的时间点(原时间 + 5次10分钟增量) df["timestamp"] = df["timestamp"].apply( lambda x: pd.date_range(start=x, periods=6, freq="10min") ) # 4. 展开时间序列列表,生成新的DataFrame df_expanded = df.explode("timestamp", ignore_index=True) # 查看结果 print(df_expanded.head(10))
效果说明:
每一行原始数据会被扩展为6行,时间依次为原小时点、+10分钟、+20分钟……+50分钟,压力、温度等数值会自动重复填充到对应的新行中。
方法二:使用resample重采样(更简洁)
如果你的数据时间序列是连续的,用Pandas的resample方法会更简洁,它会自动生成指定频率的时间点,并通过前向填充保留原始数据。
步骤代码:
import pandas as pd # 1. 构造并预处理原始DataFrame(同方法一) data = [ [1, "2016-10-01 01:00:00", 1014.7, 23.6], [2, "2016-10-01 02:00:00", 1014.3, 23.6], [3, "2016-10-01 03:00:00", 1014.3, 23.8], [4, "2016-10-01 04:00:00", 1014.3, 23.8], [5, "2016-10-01 05:00:00", 1014.4, 24.3], [6, "2016-10-01 06:00:00", 1014.9, 24.6], [7, "2016-10-01 07:00:00", 1015.6, 25.7], [8, "2016-10-01 08:00:00", 1015.8, 26], [9, "2016-10-01 09:00:00", 1016.3, 27.3], [10, "2016-10-01 10:00:00", 1016.5, 25.8], [11, "2016-10-01 11:00:00", 1016.6, 26], [12, "2016-10-01 12:00:00", 1016.6, 27.3] ] df = pd.DataFrame(data, columns=["id", "timestamp", "pressure", "temp"]) df["timestamp"] = pd.to_datetime(df["timestamp"]) # 2. 将timestamp设为索引(resample需要时间索引) df.set_index("timestamp", inplace=True) # 3. 按10分钟频率重采样,用前向填充保留上一个小时的数值 df_resampled = df.resample("10min").ffill() # 4. 重置索引,将timestamp变回普通列 df_resampled.reset_index(inplace=True) # 查看结果 print(df_resampled.head(10))
效果说明:
resample("10min")会自动生成从第一个时间点到最后一个时间点的所有10分钟间隔点,ffill()会把最近的原始数据填充到中间的时间点,最终得到的结果和方法一完全一致。
注意事项:
- 无论哪种方法,必须确保
timestamp列是datetime类型,否则无法生成时间序列或进行重采样。 - 如果你的原始数据有缺失的小时点,
resample会自动补全时间间隔,而方法一只会扩展存在的行,你可以根据自己的需求选择。
内容的提问来源于stack exchange,提问作者Nicklas Koldkjær
相关产品推荐
相关产品推荐

