行内元素数量不一致的数据集存储结构选型及随机选元素、生成0-1矩阵方案咨询
Hey there! Let's work through your problem together. You're dealing with irregularly structured data where each category has elements of different lengths, and pandas/numpy aren't fitting the bill—totally get that frustration. Here's a practical approach using Python's built-in data structures, plus code to handle all three operations you need:
First, stick to simple, native structures that play nicely with variable-length data:
- Dictionary mapping categories to element lists: Your original
datastructure is already a great starting point here—it keeps categories labeled and flexible. - Nested list: If you don't need category names, a list of lists (where each inner list holds a category's elements) works just as well.
1. Randomly Select One Element From Each Category
Use Python's random module to pick an element from each category. Tracking indices ensures accuracy even if elements are duplicated:
For Dictionary Data
import random # Your original data data = { 'func1': [[0.25, 33, 12], [12, 0, 50], [100, 25, 1]], 'func2': [[33, 2, 15], [12, 0, 50]] } # Track selected elements and their indices selected_elements = {} selected_indices = {} for category, elements in data.items(): idx = random.randint(0, len(elements) - 1) selected_indices[category] = idx selected_elements[category] = elements[idx] print("Selected elements:", selected_elements) # Example output: {'func1': [100, 25, 1], 'func2': [12, 0, 50]}
For Nested List Data
data_list = [[[0.25,33,12],[12,0,50],[100,25,1]], [[33,2,15],[12,0,50]]] selected_list = [random.choice(category) for category in data_list] print("Selected elements:", selected_list) # Example output: [[12, 0, 50], [33, 2, 15]]
2. Generate a Marker Matrix (1 for Selected, 0 Otherwise)
We'll create a structure mirroring your original data, using the stored indices to mark selected positions accurately:
marker_matrix = {} for category, elements in data.items(): # Initialize a list of 0s matching the category's length mask = [0] * len(elements) # Set the selected index to 1 mask[selected_indices[category]] = 1 marker_matrix[category] = mask print("Marker matrix:", marker_matrix) # Example output: {'func1': [0, 0, 1], 'func2': [0, 1]}
For a nested list:
marker_list = [] for i, category in enumerate(data_list): mask = [0] * len(category) mask[data_list[i].index(selected_list[i])] = 1 marker_list.append(mask) print("Marker matrix:", marker_list) # Example output: [[0, 1, 0], [1, 0]]
3. Access Specific Values of Selected Elements
This is straightforward—just index into the stored selected elements like any other list:
# Access the second value of func1's selected element print("func1 selected value (index 1):", selected_elements['func1'][1]) # Example output: 25 # Access the third value of func2's selected element print("func2 selected value (index 2):", selected_elements['func2'][2]) # Example output: 50
For the nested list version:
print("First category selected value (index 0):", selected_list[0][0]) # Example output: 12
Why Not Pandas/Numpy?
- Pandas DataFrames: While you can store nested lists in a DataFrame, operations like random selection and mask generation become clunky and non-intuitive—you're fighting against the tabular structure pandas is designed for.
- Numpy Arrays: Irregular shapes force numpy to use
objectdtype arrays, which lose most of numpy's performance benefits and trigger theVisibleDeprecationWarningyou saw. They're overkill for this use case.
内容的提问来源于stack exchange,提问作者Axix

