在类方法中为Pandas DataFrame列应用类方法时遇参数缺失错误
问题:调用类方法时触发TypeError缺失参数错误
用户创建了如下Foo类,调用实例的get_dictionary方法时触发TypeError: tweets_tokenizer() missing 1 required positional argument: 'text'错误,错误直接关联代码行df['tokenized_comments'] = df.iloc[:, c_column].apply(Foo.tweets_tokenizer)。由于tweets_tokenizer需要依赖实例的language属性,无法设置为静态方法。
代码示例
Foo类代码
class Foo: def __init__(self, file_path: str, language = None): self.language = 'italian' if language is None else language self.file_path = file_path self.file_type = file_path[-3:] def tweets_tokenizer(self, text): language = data_manager # 此处代码无实际作用 txt = word_tokenize(txt, language=self.language) # 存在参数错误 return txt def get_dictionary(self): df = self.load() # 未展示的类加载方法 c_column = int(input(f'What is the index of the column containing the comments?')) comments = df.iloc[:, c_column] df['tokenized_comments'] = df.iloc[:, c_column].apply(Foo.tweets_tokenizer) output = df.to_dict('index') return output
调用代码
item = Foo('filepath') d = item.get_dictionary()
触发错误
TypeError: tweets_tokenizer() missing 1 required positional argument: 'text'
问题原因与解决办法
原因
直接使用Foo.tweets_tokenizer调用的是未绑定实例的类方法,此时方法的第一个参数self需要手动传入,但apply只会将每行的文本作为参数传递,导致方法缺少self参数,进而报错提示缺失text(因为apply传的文本被当作了self,而text参数没被传入)。
解决办法
使用绑定实例的方法调用:将
Foo.tweets_tokenizer改为self.tweets_tokenizer,此时方法已经和当前实例绑定,apply会自动将每行文本作为text参数传入,self由实例自动提供:df['tokenized_comments'] = df.iloc[:, c_column].apply(self.tweets_tokenizer)修正
tweets_tokenizer方法内的错误:方法内存在两处无效/错误代码,需要同步修正:def tweets_tokenizer(self, text): # 移除无意义的language = data_manager代码 txt = word_tokenize(text, language=self.language) # 将txt改为text return txt
内容的提问来源于stack exchange,提问作者corvusMidnight
相关产品推荐
相关产品推荐

