You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于NumPy查找一维数组中的连续倍数集合与孤立值

Efficiently Detect Continuous 10x Multiples and Isolated Values in a Sorted NumPy Array

Great question! Given your array is sorted, contains unique elements, and all values are multiples of 10, we can leverage NumPy's vectorized operations to skip slow Python loops—this is critical for handling extremely large datasets efficiently. Here's a practical, optimized solution:

Core Idea

Since every element is a multiple of 10, we first scale the array by dividing each value by 10. This turns our problem into finding consecutive integer sequences (a continuous run of 10x multiples becomes consecutive integers after scaling). Isolated values will appear as single integers with gaps before/after them.

Step-by-Step Implementation

import numpy as np

# Example input array (replace with your large dataset)
arr = np.array([0, 10, 20, 60, 80, 90, 100])

# Step 1: Scale array to convert 10x multiples to integers (avoids floating points)
scaled_arr = arr // 10

# Step 2: Calculate differences between consecutive scaled elements
diff = np.diff(scaled_arr)

# Step 3: Find indices where the sequence breaks (difference != 1)
# Add 1 because np.diff compares index i and i+1—so the break occurs at i+1
break_indices = np.where(diff != 1)[0] + 1

# Step 4: Create full segment boundaries (include start and end of array)
segment_boundaries = np.concatenate(([0], break_indices, [len(arr)]))

# Step 5: Process each segment to separate continuous sets and isolated values
continuous_sets = []
isolated_values = []

for i in range(len(segment_boundaries) - 1):
    start_idx = segment_boundaries[i]
    end_idx = segment_boundaries[i+1]
    segment_length = end_idx - start_idx
    
    if segment_length == 1:
        # Store isolated value's index and value
        isolated_values.append({
            "index": start_idx,
            "value": arr[start_idx]
        })
    else:
        # Store continuous set's start/end indices and values
        continuous_sets.append({
            "start_index": start_idx,
            "end_index": end_idx - 1,  # end_idx is exclusive, so subtract 1 for last element index
            "start_value": arr[start_idx],
            "end_value": arr[end_idx - 1]
        })

# Print results
print("Continuous 10x multiple sets:")
for s in continuous_sets:
    print(f"Indices {s['start_index']}-{s['end_index']}: Values {s['start_value']}-{s['end_value']}")

print("\nIsolated values:")
for val in isolated_values:
    print(f"Index {val['index']}: Value {val['value']}")

Example Output

Continuous 10x multiple sets:
Indices 0-2: Values 0-20
Indices 4-6: Values 80-100

Isolated values:
Index 3: Value 60

Why This Works for Large Datasets

  • Vectorized Operations: All core steps (scaling, difference calculation, breakpoint detection) use NumPy's optimized C-backed functions—far faster than Python loops for huge arrays.
  • Minimal Looping: The only loop runs over segment boundaries, which will be drastically fewer than the total number of elements in your dataset.
  • Edge Case Handling: Works seamlessly for arrays starting/ending with isolated values, fully continuous arrays, or arrays with all isolated elements.

Quick Notes

  • Double-check that your input array is sorted and unique (you mentioned this is the case, but if not, preprocess with np.sort(np.unique(arr)) first).
  • Integer division (// 10) avoids floating-point precision issues with extremely large numbers.

内容的提问来源于stack exchange,提问作者Will

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:46:31