如何在DataFrame的map方法中无Lambda调用带参数的自定义函数?
解决方法
不用lambda的话,有几种可行方案:
1. 使用functools.partial固定参数
partial可以预先绑定函数的部分参数,生成一个新的可调用对象,刚好适配map的需求:
from functools import partial # 预先绑定percentage为0.8,生成新函数 summarize_80 = partial(summarize, percentage=.8) # 直接用map调用新函数 df['summary'] = df['text'].map(summarize_80)
2. 修改原函数为闭包形式
调整summarize函数,让它在只传入percentage时返回一个等待text参数的闭包,这样就能直接用map(summarize(percentage=.8)):
# 建议把import放在函数外部,避免重复导入影响效率 import nltk def summarize(text=None, percentage=.6): if text is None: # 返回闭包,等待传入text def process_text(txt): sentences = nltk.sent_tokenize(txt) sentences = sentences[:int(percentage*len(sentences))] return ''.join(sentences) return process_text # 正常传入text时的执行逻辑 sentences = nltk.sent_tokenize(text) sentences = sentences[:int(percentage*len(sentences))] return ''.join(sentences)
调用方式:
df['summary'] = df['text'].map(summarize(percentage=.8))
3. 使用apply替代map并传递参数
pandas的apply方法支持直接传递额外关键字参数,不需要lambda:
df['summary'] = df['text'].apply(summarize, percentage=.8)
内容的提问来源于stack exchange,提问作者Mus
相关产品推荐
相关产品推荐

