基于NumPy查找一维数组中的连续倍数集合与孤立值
Efficiently Detect Continuous 10x Multiples and Isolated Values in a Sorted NumPy Array
Great question! Given your array is sorted, contains unique elements, and all values are multiples of 10, we can leverage NumPy's vectorized operations to skip slow Python loops—this is critical for handling extremely large datasets efficiently. Here's a practical, optimized solution:
Core Idea
Since every element is a multiple of 10, we first scale the array by dividing each value by 10. This turns our problem into finding consecutive integer sequences (a continuous run of 10x multiples becomes consecutive integers after scaling). Isolated values will appear as single integers with gaps before/after them.
Step-by-Step Implementation
import numpy as np # Example input array (replace with your large dataset) arr = np.array([0, 10, 20, 60, 80, 90, 100]) # Step 1: Scale array to convert 10x multiples to integers (avoids floating points) scaled_arr = arr // 10 # Step 2: Calculate differences between consecutive scaled elements diff = np.diff(scaled_arr) # Step 3: Find indices where the sequence breaks (difference != 1) # Add 1 because np.diff compares index i and i+1—so the break occurs at i+1 break_indices = np.where(diff != 1)[0] + 1 # Step 4: Create full segment boundaries (include start and end of array) segment_boundaries = np.concatenate(([0], break_indices, [len(arr)])) # Step 5: Process each segment to separate continuous sets and isolated values continuous_sets = [] isolated_values = [] for i in range(len(segment_boundaries) - 1): start_idx = segment_boundaries[i] end_idx = segment_boundaries[i+1] segment_length = end_idx - start_idx if segment_length == 1: # Store isolated value's index and value isolated_values.append({ "index": start_idx, "value": arr[start_idx] }) else: # Store continuous set's start/end indices and values continuous_sets.append({ "start_index": start_idx, "end_index": end_idx - 1, # end_idx is exclusive, so subtract 1 for last element index "start_value": arr[start_idx], "end_value": arr[end_idx - 1] }) # Print results print("Continuous 10x multiple sets:") for s in continuous_sets: print(f"Indices {s['start_index']}-{s['end_index']}: Values {s['start_value']}-{s['end_value']}") print("\nIsolated values:") for val in isolated_values: print(f"Index {val['index']}: Value {val['value']}")
Example Output
Continuous 10x multiple sets: Indices 0-2: Values 0-20 Indices 4-6: Values 80-100 Isolated values: Index 3: Value 60
Why This Works for Large Datasets
- Vectorized Operations: All core steps (scaling, difference calculation, breakpoint detection) use NumPy's optimized C-backed functions—far faster than Python loops for huge arrays.
- Minimal Looping: The only loop runs over segment boundaries, which will be drastically fewer than the total number of elements in your dataset.
- Edge Case Handling: Works seamlessly for arrays starting/ending with isolated values, fully continuous arrays, or arrays with all isolated elements.
Quick Notes
- Double-check that your input array is sorted and unique (you mentioned this is the case, but if not, preprocess with
np.sort(np.unique(arr))first). - Integer division (
// 10) avoids floating-point precision issues with extremely large numbers.
内容的提问来源于stack exchange,提问作者Will
相关产品推荐
相关产品推荐

