You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将含for/while循环的DataFrame处理逻辑转为apply(lambda)实现

需求:将Dataframe循环处理逻辑转为apply实现

希望把以下这类嵌套循环结构:

for ... in dataframe:
   while ...:
      if ...:
         do smth
   if ...:
      do smth

转换成使用dataframe.apply(lambda ...: ...)的实现方式。

带for/while循环的示例函数

以下是实际的循环处理代码:

d_test = {
    'name' : ['South Beach', 'Dog', 'Bird', 'Ant', 'Big Dog', 'Beach', 'Dear', 'Cat', 'Fish', 'Dry Fish'],
    'cluster_number' : [1, 2, 3, 3, 2, 1, 4, 2, 2, 2]
}

df_test = pd.DataFrame(d_test)

from rapidfuzz import fuzz

df_test = df_test.sort_values(['cluster_number', 'name'])
df_test.reset_index(drop=True, inplace=True)
df_test['id'] = 0

def loop_in_cluster(index, row, df_test, index_, row_, is_i_used, i):
    while index_ < len(df_test) and df_test.loc[index, 'cluster_number'] == df_test.loc[index_, 'cluster_number'] and df_test.loc[index_, 'id'] == 0:     
        if row['name'] == df_test.loc[index_, 'name'] or fuzz.ratio(row['name'], df_test.loc[index_, 'name']) > 50:
            df_test.loc[index_,'id'] = i
            is_i_used = True
        index_ += 1
    return df_test, is_i_used
    
i = 1
is_i_used = False
for index, row in df_test.iterrows():
    row_ = row
    index_ = index
    df_test, is_i_used = loop_in_cluster(index, row, df_test, index_, row_, is_i_used, i)
    if is_i_used == True:
        i += 1
        is_i_used = False

尝试的apply实现及问题

我尝试用dataframe.apply()改写的代码如下:

i = 1
df_test.apply(lambda row: loop_in_cluster(i=i+1, index=row.name, row=row, df_test=df_test, index_=index, row_ = row, is_i_used=False) if is_i_used==True else loop_in_cluster(i=i, index=row.name, row=row, df_test=df_test, index_= index, row_=row, is_i_used=True), axis=1)

但运行时触发了StopIteration错误。我也试过用pandas的groupby.GroupBy方法,但还是觉得apply更符合我的需求。

内容的提问来源于stack exchange,提问作者illuminato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 20:01:44