如何基于两个打乱的numpy二维数组生成指定规则的新数组c?
I've worked out a solution that properly handles duplicate elements while meeting all your requirements. Let's break it down step by step.
Step-by-Step Explanation
The core issue with your previous attempt was likely checking if an element "exists" in c directly—this skips all instances of a duplicate element once it's been added once. Instead, we need to track how many times each element has been used to ensure we don't miss any duplicate instances that should be included in c.
1. Initialize c with the first i elements of a
We start by copying the first i elements of a into c—this satisfies your first requirement perfectly.
2. Track element counts
Since 2D numpy sub-arrays aren't hashable, we convert each sub-array to a tuple to use as keys in counters. We need two counters:
total_counts: The number of times each element appears ina(and thus inb, sincebis a shuffled version ofa)used_counts: The number of times each element has already been added toc
3. Populate the rest of c from b
We iterate through each element in b in order. For each element, if we haven't used up all its allowed instances (based on total_counts), we add it to c and update used_counts. We stop once c reaches the same length as a.
Full Code Implementation
import numpy as np from collections import Counter def generate_c(a, b, i): # Validate input for i if not 0 <= i <= len(a): raise ValueError("i must be between 0 and the length of array a") # Initialize c with the first i elements of a c = a[:i].copy() # If we've already filled c, return early if len(c) == len(a): return c # Convert 2D array elements to tuples for hashing in counters a_tuples = list(map(tuple, a)) b_tuples = list(map(tuple, b)) # Count total occurrences of each element in a total_counts = Counter(a_tuples) # Count how many times each element is already used in c used_counts = Counter(map(tuple, c)) # Iterate through b to fill the rest of c for elem_tuple, elem_arr in zip(b_tuples, b): if used_counts[elem_tuple] < total_counts[elem_tuple]: # Add the element to c c = np.vstack([c, elem_arr]) used_counts[elem_tuple] += 1 # Stop once c matches the length of a if len(c) == len(a): break return c
Test the Solution
Example from Your Question
a = np.array([[0, 1], [2, 3], [4, 5], [6, 7]]) b = np.array([[2, 3], [6, 7], [0, 1], [4, 5]]) i = 2 c = generate_c(a, b, i) print(c)
Output:
[[0 1] [2 3] [6 7] [4 5]]
Test with Duplicate Elements
# Test case with duplicates a = np.array([[0, 1], [0, 1], [2, 3]]) b = np.array([[0, 1], [2, 3], [0, 1]]) i = 1 c = generate_c(a, b, i) print(c)
Output:
[[0 1] [0 1] [2 3]]
This solution ensures c always has the same length as a and b, correctly handles duplicates, and adheres to both of your requirements.
内容的提问来源于stack exchange,提问作者Sergey Ronin

