You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Dataframe元素遍历分类至列表问题求助

Hey there! Let's work through this problem together—since you're dealing with a pretty large DataFrame (1000×3000), efficiency and correctness are key, and we’ll make sure we don’t touch the original data at all.

First, let’s break down why your two approaches ran into issues:

  • Approach 1 problem: When you use df.apply(lambda x: function103(x), axis=1), the x passed to your function is an entire row (a Series), not individual elements. Checking x > 0 on a Series returns another Series of booleans, and Python can’t use that directly in an if statement—hence the ambiguous truth value error.
  • Approach 2 problem: Using (x>0).any() checks if any element in the row is greater than 0, not each element individually. That’s why you ended up storing entire Series objects instead of single values.

For large DataFrames, we want to avoid slow element-wise loops where possible. Here are two efficient, straightforward methods:

Method 1: Vectorized Pandas Operations (Fastest & Cleanest)

Pandas is built for vectorized operations—we can flatten the DataFrame into a 1-dimensional Series, then filter values in one go:

import pandas as pd

# Assume your DataFrame is named `df`
# Flatten the DataFrame to a 1D Series (preserves all elements)
flattened_data = df.stack()

# Split into your target lists
A1 = flattened_data[flattened_data > 0].tolist()
A2 = flattened_data[flattened_data <= 0].tolist()

This method is lightning-fast because it leverages pandas' optimized backend, and it never modifies the original DataFrame. The stack() method just reshapes the data temporarily, and tolist() converts the filtered results directly into lists of individual values.

Method 2: Numpy Array Traversal (Good for Explicit Loops)

If you prefer working with a more explicit loop (easier to follow as a beginner), convert the DataFrame to a numpy array first—numpy’s array operations are far faster than looping through pandas rows/columns:

# Extract the underlying numpy array and flatten it to 1D
all_values = df.values.flatten()

A1 = []
A2 = []
for val in all_values:
    if val > 0:
        A1.append(val)
    else:
        A2.append(val)

This avoids the overhead of pandas Series objects during iteration, making it much more efficient than using iterrows() or applymap() for large datasets.

Why We Avoid applymap() for Large Data

While you could use df.applymap() to run a function on every element, it’s not recommended for 3 million elements (1000×3000). applymap() is essentially a loop under the hood, and it will be significantly slower than the vectorized methods above.

内容的提问来源于stack exchange,提问作者Max Bauer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:45:48