You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何映射数组子索引至原索引以实现时间序列参数按频率分组

解决按频率分组并获取原数组索引的问题

看起来你已经把大部分基础组件都搭建好了,只差最后一步的整合!我来帮你梳理两种可行的方案:一种是基于你现有的代码组件进行整合,另一种是更简洁的直接实现方式,你可以根据自己的需求选择。

方案一:基于你现有组件的整合

你已经有了按频率排序的sorted_data子索引、原索引映射字典,只需要把这些组件串联起来,再处理同频率组的排序即可:

import numpy as np

# 你已有的变量定义和函数(保留关键部分)
speed = np.array([4, 6, 8, 3, 6, 9, 7, 6, 4, 3])*100
elap_hr = sorted(np.random.randint(low=1, high=40, size=10))

def get_sorted_data(data, sort_type='default'):
    if sort_type == 'default':
        res = sorted(data)
    elif sort_type == 'argsort':
        res = np.argsort(data)
    elif sort_type == 'by size':
        res = sorted(data, key=len)
    return res

def sort_data_by_frequency(data):
    uniq_data = np.unique(data)
    sorted_data = get_sorted_data(data)
    res = [np.where(sorted_data == uniq_data[i])[0] for i in range(len(uniq_data))]
    res = get_sorted_data(res, 'by size')[::-1]
    return res

def get_dictionary_mapping(keys, values):
    return dict(zip(keys, values))

# 你已有的计算结果
sorted_speed = get_sorted_data(speed)
argsorted_speed = get_sorted_data(speed, 'argsort')
freqsorted_speed = sort_data_by_frequency(speed)
idx_orig = np.array([i for i in range(len(argsorted_speed))], dtype=int)
index_to_index_map = get_dictionary_mapping(idx_orig, argsorted_speed)

# ------------------- 整合步骤开始 -------------------
# 1. 将频率组转换为原索引,并记录对应数值
groups_with_values = []
for sorted_sub_idx in freqsorted_speed:
    # 同一组内的数值完全相同,取第一个元素对应的值即可
    value = sorted_speed[sorted_sub_idx[0]]
    # 通过映射字典转换为原数组索引
    original_indices = [index_to_index_map[idx] for idx in sorted_sub_idx]
    groups_with_values.append( (value, original_indices) )

# 2. 按「频率降序,同频率按数值升序」排序这些组
groups_with_values.sort(key=lambda x: (-len(x[1]), x[0]))

# 3. 提取最终的原索引分组结果
final_result = [group[1] for group in groups_with_values]
print(final_result)
# 输出:[[1, 4, 7], [3, 9], [0, 8], [6], [2], [5]]

说明:

  • 第一步把你得到的sorted_data子索引,通过index_to_index_map转换成原数组的索引,同时记录每个组对应的数值,方便后续排序。
  • 第二步用排序键(-len(x[1]), x[0])实现:先按组的长度(频率)降序,长度相同则按数值升序,完美匹配你的需求。

方案二:更简洁的直接实现(推荐)

如果不需要保留之前的中间组件,直接用numpy的内置功能可以更高效地完成需求,代码更简洁:

import numpy as np

speed = np.array([4, 6, 8, 3, 6, 9, 7, 6, 4, 3])*100

# 1. 获取唯一值和对应的出现频率
uniq_values, counts = np.unique(speed, return_counts=True)

# 2. 按「频率降序,同频率按数值升序」排序唯一值
# np.lexsort会先按后面的参数排序,这里先按-counts(频率降序),再按uniq_values(数值升序)
sorted_indices = np.lexsort((uniq_values, -counts))
sorted_unique = uniq_values[sorted_indices]

# 3. 对每个排序后的唯一值,提取原数组中所有匹配的索引
final_result = [np.where(speed == val)[0].tolist() for val in sorted_unique]
print(final_result)
# 输出:[[1, 4, 7], [3, 9], [0, 8], [6], [2], [5]]

说明:

  • np.unique(..., return_counts=True)可以一次性拿到所有唯一值和它们的出现次数,省去手动统计的麻烦。
  • np.lexsort是numpy的多键排序工具,非常适合这种需要复合排序规则的场景。
  • 最后直接用np.where提取每个值对应的原索引,一步到位。

内容的提问来源于stack exchange,提问作者user7345804

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:42:35