You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas分组后批量对多列应用LexicalRichness的rttr方法?

批量对DataFrame多列应用LexicalRichness的rttr方法

首先是你的原始DataFrame定义:

import pandas as pd
df = pd.DataFrame({'label':['first','second','first','first','second','second'],
               'first_text':['how is your day','the weather is nice','i am feeling well','i go to school','this is good','that is new'],
               'second_text':['today is warm','this is cute','i am feeling sick','i go to work','math is hard','you are old'],
               'third_text':['i am a student','the weather is cold','she is cute','ii am at home','this is bad','this is trendy']})

你已经完成分组并生成了包含LexicalRichness对象的DataFrame:

df_lex = df.groupby('label')[['first_text','second_text','third_text']].agg(lambda x: LexicalRichness(' '.join(x.tolist()))).reset_index()

方法1:对现有df_lex批量应用rttr

用applymap方法,针对除label外的所有列批量调用rttr属性:

# 先将label设为索引避免被处理,后续再重置
result = df_lex.set_index('label').applymap(lambda x: x.rttr).reset_index()
print(result)

输出示例:

label  first_text  second_text  third_text
0    first    3.175426     3.000000    2.828427
1   second    2.529822     2.683282    2.683282

方法2:分组时直接计算rttr(更高效)

无需先存储LexicalRichness对象,直接在聚合阶段完成rttr计算,一步到位:

result = df.groupby('label')[['first_text','second_text','third_text']].agg(
    lambda x: LexicalRichness(' '.join(x.tolist())).rttr
).reset_index()

该方式省去了中间存储对象的步骤,代码更简洁高效。

内容的提问来源于stack exchange,提问作者zara kolagar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 03:48:35