You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并两个NumPy数组并去重,优先保留第一个数组元素?

Hey there! Let's break down how to solve this NumPy array problem, and figure out which approach makes more sense for your needs.

First, let's restate your core requirement clearly: keep all elements from array1 in their original order, then add only the elements from array2 that don't already exist in array1 (also keeping their order from array2, no duplicates added). The final result should be [1,6,7,9,3,5,8,2].

This method directly targets your needs—we only add the elements from array2 that array1 doesn't have, avoiding any unnecessary duplicate handling.

Here's the code:

import numpy as np

array1 = np.array([1,6,7,9,3,5])
array2 = np.array([3,5,8,9,2])

# Get elements from array2 that aren't present in array1
unique_to_array2 = array2[~np.isin(array2, array1)]
# Combine array1 with these unique elements
final_array = np.concatenate([array1, unique_to_array2])

print(final_array)
# Output: [1 6 7 9 3 5 8 2]

Why this works:

  • np.isin(array2, array1) checks each element in array2 to see if it exists in array1, returning a boolean array. The ~ inverts this to get elements not in array1.
  • This is a vectorized operation, so it's way faster than looping through each element in Python (especially for large arrays).
  • Most importantly: it preserves every element from array1, even if array1 has duplicates (if that's ever a case in your data). You're only adding the exact elements you need from array2.

Approach 2: Merge First, Then Deduplicate

If you want to try merging first then removing duplicates, you have to be careful—regular np.unique() will sort the result, which breaks your desired order. Instead, we can use np.unique with return_index=True to keep the original order:

import numpy as np

array1 = np.array([1,6,7,9,3,5])
array2 = np.array([3,5,8,9,2])

# Merge the two arrays first
combined_array = np.concatenate([array1, array2])
# Get unique elements while preserving their first occurrence order
_, unique_indices = np.unique(combined_array, return_index=True)
final_array = combined_array[np.sort(unique_indices)]

print(final_array)
# Output: [1 6 7 9 3 5 8 2]

Caveats with this approach:

  • If array1 has duplicate elements (e.g., array1 = [1,6,7,9,3,5,3]), this method will remove those duplicates from array1, which violates your goal of "keeping as many elements from the first array as possible".
  • It uses more memory because you first create a larger combined array before processing. For small arrays this isn't a big deal, but it's inefficient for large datasets.

Which One Is Better?

Go with Approach 1 for almost all cases:

  • It's more efficient (both in memory and speed).
  • It strictly adheres to your requirement of preserving all elements from array1.
  • It's more readable—anyone looking at the code can immediately tell you're adding only new elements from array2.

Approach 2 only makes sense if you specifically need to handle duplicates within array1 (but your question says you want to keep array1's elements, so this is unlikely).

内容的提问来源于stack exchange,提问作者Will.S89

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:34:16