如何基于interval_df的深度区间填充sample_df的对应列?
基于深度区间填充离散采样数据
需求:给包含离散采样深度的sample_df填充目标列,逻辑为遍历存储深度区间及对应值的interval_df,判断sample_df中每个深度所属的区间,将该区间的value值赋值给sample_df的对应行。
示例数据
区间数据(interval_df)
import pandas as pd interval_df = pd.DataFrame({ 'top depth':[100,200,700], 'bottom depth':[200,700,1000], 'value':[15,10,20], })
采样数据(sample_df)
sample_df = pd.DataFrame({ 'depth':[258,300,567,858,900], 'value':[0,0,0,0,0] })
解决方案
方法1:遍历匹配(直观易懂)
适合数据量较小的场景,逻辑清晰易调试:
# 初始化目标列 sample_df['uncertainty'] = 0 # 逐个匹配采样深度对应的区间 for idx, depth in enumerate(sample_df['depth']): # 筛选当前深度所属的区间 match_interval = interval_df[(interval_df['top depth'] <= depth) & (depth < interval_df['bottom depth'])] # 赋值对应区间的value sample_df.loc[idx, 'uncertainty'] = match_interval['value'].values[0]
方法2:矢量化操作(高效处理大数据)
利用pandas内置函数实现批量匹配,性能更优:
# 整理区间边界和对应标签值 bins = interval_df['top depth'].tolist() + [interval_df['bottom depth'].iloc[-1]] labels = interval_df['value'].tolist() # 批量匹配区间并赋值 sample_df['uncertainty'] = pd.cut( sample_df['depth'], bins=bins, labels=labels, include_lowest=True # 确保左边界包含在内 )
期望输出
sample_df = pd.DataFrame({ 'depth':[258,300,567,858,900], 'uncertainty':[10,10,10,20,20] })
内容的提问来源于stack exchange,提问作者Eirik Gottschalk Ballo
相关产品推荐
相关产品推荐

