You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从DataFrame提取数据并使用apply方法填充分组列表

Solution: Populate Lists from DataFrame Rows Using apply()

Got it, let's break this down step by step. I'll use a sample DataFrame that mirrors the structure you described, then show you how to use row-wise apply() to build your target list of lists.

Step 1: Set Up Sample Data

First, let's define a sample DataFrame to work with (replace this with your actual data):

import pandas as pd

df = pd.DataFrame({
    'a': [1, 1, 2, 2, 2],
    'b': [10, 20, 30, 40, 50]
})

Here, column a has 2 unique values (1 and 2), so our final result will be a list with 2 sublists—one for each unique value in a, filled with corresponding b values.

Step 2: Initialize Your Result Structure

Before using apply(), we need to set up a container for our results and a quick way to map values in a to their position in the result list:

# Get unique values from 'a' (keep order as they appear; use sorted() if you need sorted order)
unique_a_values = df['a'].unique().tolist()
# Create a list of empty sublists, one per unique value in 'a'
result_lists = [[] for _ in unique_a_values]
# Make a dictionary to map each 'a' value to its index in result_lists (O(1) lookups)
a_value_to_index = {val: idx for idx, val in enumerate(unique_a_values)}

This dictionary is key—it keeps the row-wise processing efficient, even for large DataFrames, since dictionary lookups are lightning fast.

Step 3: Use apply() to Fill the Lists

Next, define a function that takes each row, finds the correct sublist, and appends the b value to it. Then apply this function to every row:

def populate_sublist(row):
    # Get the index of the sublist matching the current row's 'a' value
    sublist_index = a_value_to_index[row['a']]
    # Append the row's 'b' value to the corresponding sublist
    result_lists[sublist_index].append(row['b'])

# Apply the function row-by-row (axis=1 specifies we're working with rows)
df.apply(populate_sublist, axis=1)

Step 4: Check the Result

After running the apply() call, your result_lists will be filled exactly as you need:

print(result_lists)
# Output: [[10, 20], [30, 40, 50]]

Notes for Large DataFrames

  • The dictionary lookup ensures this method stays efficient even with millions of rows—no slow searching through lists for each row.
  • If you need the unique values in a sorted instead of their original order, replace df['a'].unique().tolist() with sorted(df['a'].unique()).
  • Quick side note: For extremely large datasets, df.groupby('a')['b'].apply(list) might be faster than row-wise apply(), but since you specifically requested using apply() on rows, the above method fits your requirement perfectly.

内容的提问来源于stack exchange,提问作者ALejandro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:02:20