Python 2.7中ndarray删除行列的更优实现方法问询
Hey there! Let's break this down for you—first we'll clear up why a[rows_to_keep, columns_to_keep] gave you unexpected results, then share cleaner, more efficient ways to delete rows and columns in NumPy.
Why a[rows_to_keep, columns_to_keep] Didn't Work as Expected
NumPy's indexing rules are a bit counterintuitive here. When you pass two 1D arrays directly into the index (like your rows_to_keep and columns_to_keep), it doesn't slice rows first then columns. Instead, it selects individual elements at matching (row, column) pairs.
For example, if rows_to_keep = [0, 2] and columns_to_keep = [1, 2], a[[0,2], [1,2]] would pull elements at positions (0,1) and (2,2)—resulting in a 1D array [1, 8], not the 2x2 matrix you wanted.
Your original approach a[rows_to_keep,:][:,columns_to_keep] works because it’s two separate operations: first slice all columns for your target rows, then slice all those rows for your target columns—this preserves the 2D structure you need.
Cleaner, More Efficient Implementations
1. Use np.delete (Recommended)
NumPy has a built-in np.delete function made exactly for this task. It’s readable, concise, and avoids manually creating "keep" indices:
import numpy as np number_of_rows = 3 number_of_columns = 3 a = np.arange(number_of_rows*number_of_columns).reshape(number_of_rows, number_of_columns) rows_to_remove = [1] columns_to_remove = [0] # Delete rows first, then columns result = np.delete(np.delete(a, rows_to_remove, axis=0), columns_to_remove, axis=1) print(result) # Output: # [[1 2] # [7 8]]
axis=0targets rows,axis=1targets columns- You can pass single indices or lists to delete multiple rows/columns at once
2. Improve Your "Keep" Index Method (If You Prefer This Approach)
If you want to stick with generating "keep" indices, avoid using set—sets are unordered, so you might end up with rows/columns in the wrong order (especially in Python 2.7). Use a list comprehension instead to preserve sequence:
rows_to_keep = [i for i in range(a.shape[0]) if i not in rows_to_remove] columns_to_keep = [i for i in range(a.shape[1]) if i not in columns_to_remove] result = a[rows_to_keep, :][:, columns_to_keep]
This ensures your rows/columns stay in their original order, which is critical if sequence matters for your data.
Final Takeaway
np.delete is almost always the best choice here—it leverages NumPy's optimized operations, is easy to read, and eliminates the need for manual index wrangling. No need to reinvent the wheel unless you have a specific use case for the "keep" index approach.
内容的提问来源于stack exchange,提问作者Rafael Valero

