Python时间序列间歇性信号分类代码优化及专业实现方案咨询
Python时间序列间歇性信号分类代码优化及专业实现方案咨询
问题背景
我需要处理传感器产生的间歇性信号:信号会在一个周期为0.01,下一个周期为0,再下一个周期又回到0.01,这种是设计内的正常情况。我的目标是忽略最多N个周期的间隙,检测出有效信号(允许非实时分析,也就是可以向前查看后续数据)。举个例子,如果允许忽略最多2个周期的间隙,检测结果如下:
| 信号值 | 检测结果 |
|---|---|
| 0 | FALSE |
| 0.01 | TRUE |
| 0.036 | TRUE |
| 0 | TRUE |
| 0.2 | TRUE |
| 0 | FALSE |
| 0 | FALSE |
| 0 | FALSE |
| 0 | FALSE |
| 0.5 | TRUE |
| 0 | TRUE |
| 0 | TRUE |
| 0.1 | TRUE |
| 0.0 | FALSE |
| 0.0 | FALSE |
| 0.0 | FALSE |
我自己写了一个初级版本的函数,但它的缺陷是:只会把检测结果按忽略的间隙长度延长,不会向前查看间隙内是否后续还有有效信号。下面是我写的代码:
from IPython.display import display import pandas as pd def find_continuous(df, threshold, max_gap): # df - 包含数据的序列 # threshold - 检测的最小值(包含该值) # max_gap - 最后一个>=threshold的值之后,仍被视为有效信号的最大周期数 i = 0 min_value = threshold currentlyPriming = False primeTimes = [] PrimeTrue = [] Prime2 = [] distance = 0 distance_to_check = 0 distance_checked = 1 while i < (len(df)): print('element equals ', df.iloc[i], ', index is ', df.index[i], ', current i is ', i) if df.iloc[i] < min_value and len(PrimeTrue) > 0: currentlyPriming = False print ('currently priming set to False') print('last index element in primeTrue list is ', PrimeTrue[-1]) if max_gap == 0: print('max gap is at zero') elif max_gap == 1 and df.index[i] - PrimeTrue[-1] == 1: print('max gap is 1 and this element is next after positive') primeTimes.append(df.index[i]) elif max_gap >= 2: try: distance = (Prime2[-1] - PrimeTrue[-1]) distance_to_check = max(max_gap - distance, 0) print('last index element in Prime2 list is ', Prime2[-1]) except: print('Prime2 has not been initiated, first clustering detection') distance = 88888 distance_to_check = max(max_gap - 1, 0) print('distance is ', distance, ' distance to check is ', distance_to_check ) if distance_to_check > 0: primeTimes.append(df.index[i]) Prime2.append(df.index[i]) distance_checked += 1 print('distance checked is ', distance_checked) elif distance_to_check == 0: distance_checked = 1 elif df.iloc[i] < min_value and len(PrimeTrue) == 0: currentlyPriming = False print('element is less than minimum value and element greater than minimum value was not found yet') elif df.iloc[i] >= min_value: PrimeTrue.append(df.index[i]) if currentlyPriming: primeTimes.append(df.index[i]) print('section d, priming is ', currentlyPriming ) elif not currentlyPriming: primeTimes.append(df.index[i]) currentlyPriming = True print('section f, priming is ', currentlyPriming ) i += 1 return primeTimes if __name__ == "__main__": values = [0.05,0,0,0,0,0.037037037,0,0,0,0.035714286,0,0.05,0,0,0,0,0,0,0,0.025677,0,0.05,0,0,0,0.04,0,0.031037037,0,0,0,0,0,0.04,0,0,0,0.074074074,0,0.032258065,0,0,0,0.001,0,0,0,0,0,0,0,0,0,0,0.060606061,0,0,0,0.060606061,0,0,0,0,0,0,0,0,0] v1 = pd.DataFrame(data=values, index=None, columns=['values']) list2 = [] list2 = find_continuous(v1['values'], 0.035, 2) for k in range(len(list2)): print(k) v1.at[list2[k],'cluster'] = list2[k] with pd.option_context("display.max_rows", v1.shape[0]): display(v1)
我的疑问
有没有更好的Python实现方式?专业的Python开发者会怎么写这段代码?
备注:内容来源于stack exchange,提问作者Djangu
相关产品推荐
相关产品推荐

