如何映射数组子索引至原索引以实现时间序列参数按频率分组
解决按频率分组并获取原数组索引的问题
看起来你已经把大部分基础组件都搭建好了,只差最后一步的整合!我来帮你梳理两种可行的方案:一种是基于你现有的代码组件进行整合,另一种是更简洁的直接实现方式,你可以根据自己的需求选择。
方案一:基于你现有组件的整合
你已经有了按频率排序的sorted_data子索引、原索引映射字典,只需要把这些组件串联起来,再处理同频率组的排序即可:
import numpy as np # 你已有的变量定义和函数(保留关键部分) speed = np.array([4, 6, 8, 3, 6, 9, 7, 6, 4, 3])*100 elap_hr = sorted(np.random.randint(low=1, high=40, size=10)) def get_sorted_data(data, sort_type='default'): if sort_type == 'default': res = sorted(data) elif sort_type == 'argsort': res = np.argsort(data) elif sort_type == 'by size': res = sorted(data, key=len) return res def sort_data_by_frequency(data): uniq_data = np.unique(data) sorted_data = get_sorted_data(data) res = [np.where(sorted_data == uniq_data[i])[0] for i in range(len(uniq_data))] res = get_sorted_data(res, 'by size')[::-1] return res def get_dictionary_mapping(keys, values): return dict(zip(keys, values)) # 你已有的计算结果 sorted_speed = get_sorted_data(speed) argsorted_speed = get_sorted_data(speed, 'argsort') freqsorted_speed = sort_data_by_frequency(speed) idx_orig = np.array([i for i in range(len(argsorted_speed))], dtype=int) index_to_index_map = get_dictionary_mapping(idx_orig, argsorted_speed) # ------------------- 整合步骤开始 ------------------- # 1. 将频率组转换为原索引,并记录对应数值 groups_with_values = [] for sorted_sub_idx in freqsorted_speed: # 同一组内的数值完全相同,取第一个元素对应的值即可 value = sorted_speed[sorted_sub_idx[0]] # 通过映射字典转换为原数组索引 original_indices = [index_to_index_map[idx] for idx in sorted_sub_idx] groups_with_values.append( (value, original_indices) ) # 2. 按「频率降序,同频率按数值升序」排序这些组 groups_with_values.sort(key=lambda x: (-len(x[1]), x[0])) # 3. 提取最终的原索引分组结果 final_result = [group[1] for group in groups_with_values] print(final_result) # 输出:[[1, 4, 7], [3, 9], [0, 8], [6], [2], [5]]
说明:
- 第一步把你得到的
sorted_data子索引,通过index_to_index_map转换成原数组的索引,同时记录每个组对应的数值,方便后续排序。 - 第二步用排序键
(-len(x[1]), x[0])实现:先按组的长度(频率)降序,长度相同则按数值升序,完美匹配你的需求。
方案二:更简洁的直接实现(推荐)
如果不需要保留之前的中间组件,直接用numpy的内置功能可以更高效地完成需求,代码更简洁:
import numpy as np speed = np.array([4, 6, 8, 3, 6, 9, 7, 6, 4, 3])*100 # 1. 获取唯一值和对应的出现频率 uniq_values, counts = np.unique(speed, return_counts=True) # 2. 按「频率降序,同频率按数值升序」排序唯一值 # np.lexsort会先按后面的参数排序,这里先按-counts(频率降序),再按uniq_values(数值升序) sorted_indices = np.lexsort((uniq_values, -counts)) sorted_unique = uniq_values[sorted_indices] # 3. 对每个排序后的唯一值,提取原数组中所有匹配的索引 final_result = [np.where(speed == val)[0].tolist() for val in sorted_unique] print(final_result) # 输出:[[1, 4, 7], [3, 9], [0, 8], [6], [2], [5]]
说明:
np.unique(..., return_counts=True)可以一次性拿到所有唯一值和它们的出现次数,省去手动统计的麻烦。np.lexsort是numpy的多键排序工具,非常适合这种需要复合排序规则的场景。- 最后直接用
np.where提取每个值对应的原索引,一步到位。
内容的提问来源于stack exchange,提问作者user7345804
相关产品推荐
相关产品推荐

