Python Dataframe元素遍历分类至列表问题求助
Hey there! Let's work through this problem together—since you're dealing with a pretty large DataFrame (1000×3000), efficiency and correctness are key, and we’ll make sure we don’t touch the original data at all.
First, let’s break down why your two approaches ran into issues:
- Approach 1 problem: When you use
df.apply(lambda x: function103(x), axis=1), thexpassed to your function is an entire row (a Series), not individual elements. Checkingx > 0on a Series returns another Series of booleans, and Python can’t use that directly in anifstatement—hence the ambiguous truth value error. - Approach 2 problem: Using
(x>0).any()checks if any element in the row is greater than 0, not each element individually. That’s why you ended up storing entire Series objects instead of single values.
Recommended Solutions
For large DataFrames, we want to avoid slow element-wise loops where possible. Here are two efficient, straightforward methods:
Method 1: Vectorized Pandas Operations (Fastest & Cleanest)
Pandas is built for vectorized operations—we can flatten the DataFrame into a 1-dimensional Series, then filter values in one go:
import pandas as pd # Assume your DataFrame is named `df` # Flatten the DataFrame to a 1D Series (preserves all elements) flattened_data = df.stack() # Split into your target lists A1 = flattened_data[flattened_data > 0].tolist() A2 = flattened_data[flattened_data <= 0].tolist()
This method is lightning-fast because it leverages pandas' optimized backend, and it never modifies the original DataFrame. The stack() method just reshapes the data temporarily, and tolist() converts the filtered results directly into lists of individual values.
Method 2: Numpy Array Traversal (Good for Explicit Loops)
If you prefer working with a more explicit loop (easier to follow as a beginner), convert the DataFrame to a numpy array first—numpy’s array operations are far faster than looping through pandas rows/columns:
# Extract the underlying numpy array and flatten it to 1D all_values = df.values.flatten() A1 = [] A2 = [] for val in all_values: if val > 0: A1.append(val) else: A2.append(val)
This avoids the overhead of pandas Series objects during iteration, making it much more efficient than using iterrows() or applymap() for large datasets.
Why We Avoid applymap() for Large Data
While you could use df.applymap() to run a function on every element, it’s not recommended for 3 million elements (1000×3000). applymap() is essentially a loop under the hood, and it will be significantly slower than the vectorized methods above.
内容的提问来源于stack exchange,提问作者Max Bauer

