如何在Pandas分组后批量对多列应用LexicalRichness的rttr方法?
批量对DataFrame多列应用LexicalRichness的rttr方法
首先是你的原始DataFrame定义:
import pandas as pd df = pd.DataFrame({'label':['first','second','first','first','second','second'], 'first_text':['how is your day','the weather is nice','i am feeling well','i go to school','this is good','that is new'], 'second_text':['today is warm','this is cute','i am feeling sick','i go to work','math is hard','you are old'], 'third_text':['i am a student','the weather is cold','she is cute','ii am at home','this is bad','this is trendy']})
你已经完成分组并生成了包含LexicalRichness对象的DataFrame:
df_lex = df.groupby('label')[['first_text','second_text','third_text']].agg(lambda x: LexicalRichness(' '.join(x.tolist()))).reset_index()
方法1:对现有df_lex批量应用rttr
用applymap方法,针对除label外的所有列批量调用rttr属性:
# 先将label设为索引避免被处理,后续再重置 result = df_lex.set_index('label').applymap(lambda x: x.rttr).reset_index() print(result)
输出示例:
label first_text second_text third_text 0 first 3.175426 3.000000 2.828427 1 second 2.529822 2.683282 2.683282
方法2:分组时直接计算rttr(更高效)
无需先存储LexicalRichness对象,直接在聚合阶段完成rttr计算,一步到位:
result = df.groupby('label')[['first_text','second_text','third_text']].agg( lambda x: LexicalRichness(' '.join(x.tolist())).rttr ).reset_index()
该方式省去了中间存储对象的步骤,代码更简洁高效。
内容的提问来源于stack exchange,提问作者zara kolagar
相关产品推荐
相关产品推荐

