如何在循环中扩展NumPy数组?并将姓名添加到出勤数组最左侧
Hey there! Let's work through your two NumPy questions using the dataset you provided. I'll break each one down with practical examples that fit your specific use case.
First, a quick heads-up: NumPy arrays are designed to be fixed-size for performance. Extending them in a loop (like you might with a Python list) can be inefficient because each extension requires copying the entire array to a new memory block. But if you need to do it, here's how—and a better alternative for most cases.
Option 1: Dynamic extension in a loop (use sparingly)
You can use np.vstack() (for rows) or np.hstack() (for columns) to append new elements in each iteration. Let's simulate adding 3 new attendance rows to your existing array:
import numpy as np # Your existing attendance array attendance = np.array([[1, 0, 0, 1, 1, 1, 0], [1, 0, 1, 1, 1, 1, 1], [1, 1, 1, 1, 1, 1, 0], [1, 0, 0, 1, 1, 1, 1], [1, 1, 1, 0, 1, 1, 1], [1, 0, 0, 0, 1, 1, 1], [1, 1, 1, 1, 1, 1, 1], [1, 1, 0, 0, 0, 0, 0], [1, 1, 1, 1, 1, 1, 1], [0, 0, 0, 1, 1, 1, 1]]) # Initialize empty array with matching column count extended_arr = np.array([], dtype=int).reshape(0, 7) # Add existing rows first (or start empty if building from scratch) extended_arr = np.vstack([extended_arr, attendance]) # Loop to add new rows (simulating dynamic data) for _ in range(3): # Generate a random attendance row for example new_row = np.random.randint(0, 2, size=7) extended_arr = np.vstack([extended_arr, new_row]) print(extended_arr.shape) # Output: (13, 7) — 10 original + 3 new rows
Option 2: Pre-allocate the array (far more efficient)
If you know the final size of your array upfront, pre-allocate memory first, then fill it in the loop. This avoids repeated array copies:
# We know we want 13 total rows (10 original + 3 new) total_rows = 13 pre_allocated = np.empty((total_rows, 7), dtype=int) # Fill existing data first pre_allocated[:10] = attendance # Loop to fill new rows for i in range(10, total_rows): pre_allocated[i] = np.random.randint(0, 2, size=7) print(pre_allocated.shape) # Output: (13, 7)
This is the preferred method for large datasets or long loops—it's way faster and uses memory more efficiently.
name_list to the leftmost column of the corresponding row in attendance? Since name_list contains strings and attendance contains integers, we can't combine them directly into a standard NumPy array (which requires homogeneous data types). Instead, we have two good options:
Option 1: Use an object-type array (mixed data types)
Convert attendance to an object-type array (which can hold any Python object, including strings and integers), then stack the names as a new left column:
name_list = np.array(["Ali","Ahmad","Beng","Chris","Sita","Marion","Stephen","Cobby","Akmal","Nita"]) # Reshape names into a 2D array (1 column, 10 rows) name_col = name_list.reshape(-1, 1) # Convert attendance to object type to support mixing with strings attendance_obj = attendance.astype(object) # Horizontally stack the name column and attendance data final_array = np.hstack([name_col, attendance_obj]) # Print the result for row in final_array: print(row)
Sample output:
['Ali' 1 0 0 1 1 1 0] ['Ahmad' 1 0 1 1 1 1 1] ['Beng' 1 1 1 1 1 1 0] ...
Option 2: Use a structured array (tabular data with named columns)
If you want a more structured, table-like format (with named columns for each day and the name), use a structured array. This lets you access columns by name (e.g., result['name'] or result['Sunday']):
day_list = np.array(["Sunday","Monday","Tuesday","Wednesday","Thursday","Friday","Saturday"]) # Define the data type: name (string) + one integer column per day dtype_spec = [('name', 'U10')] + [(day, int) for day in day_list] # Create an empty structured array with the right length structured_result = np.empty(len(name_list), dtype=dtype_spec) # Fill in the names structured_result['name'] = name_list # Fill in attendance data for each day for idx, day in enumerate(day_list): structured_result[day] = attendance[:, idx] # Access data by column name print("Names:", structured_result['name']) print("Sunday attendance:", structured_result['Sunday'])
This is great if you plan to work with the data like a table, as it keeps each column's type clear and accessible.
内容的提问来源于stack exchange,提问作者Ayman hebat

