Pandas检查列是否含指定值赋值时报TypeError如何解决
问题描述
需要遍历整个DataFrame逐行检查text列是否包含指定列表中的关键词,若当前行匹配到关键词则将新增的test列对应值设为True,否则设为False。运行代码时抛出错误:TypeError: 'bool' object does not support item assignment。
示例DataFrame数据
customerId text 0 1 Something with Cat 1 3 That is a huge dog 2 3 Hello agian
原实现代码
import pandas as pd import copy import re d = { "customerId": [1, 3, 3], "text": ["Something with Cat", "That is a huge dog", "Hello agian"], } df = pd.DataFrame(data=d) my_list = ['cat', 'dog', 'mouse'] def f(x): match = False for element in my_list: x = bool(re.search(element, x['text'], re.IGNORECASE)) if(x): match = True break x['test'] = str(match) return x df['test'] = None df = df.apply(lambda x: f(x), axis = 1)
期望输出结果
customerId text test 0 1 Something with Cat True 1 3 That is a huge dog True 2 3 Hello agian False
完整错误栈
--------------------------------------------------------------------------- TypeError Traceback (most recent call last) File <timed exec>:13, in <module> File ~\Anaconda3\lib\site-packages\pandas\core\frame.py:8839, in DataFrame.apply(self, func, axis, raw, result_type, args, **kwargs) 8828 from pandas.core.apply import frame_apply 8830 op = frame_apply( 8831 self, 8832 func=func, (...) 8837 kwargs=kwargs, 8838 ) -> 8839 return op.apply().__finalize__(self, method="apply") File ~\Anaconda3\lib\site-packages\pandas\core\apply.py:727, in FrameApply.apply(self) 724 elif self.raw: 725 return self.apply_raw() --> 727 return self.apply_standard() File ~\Anaconda3\lib\site-packages\pandas\core\apply.py:851, in FrameApply.apply_standard(self) 850 def apply_standard(self): --> 851 results, res_index = self.apply_series_generator() 853 # wrap results 854 return self.wrap_results(results, res_index) File ~\Anaconda3\lib\site-packages\pandas\core\apply.py:867, in FrameApply.apply_series_generator(self) 864 with option_context("mode.chained_assignment", None): 865 for i, v in enumerate(series_gen): 866 # ignore SettingWithCopy here in case the user mutates --> 867 results[i] = self.f(v) 868 if isinstance(results[i], ABCSeries): 869 # If the return value is a view of v, make a copy to avoid errors from underlying data swap 870 results[i] = results[i].copy(deep=False) File <timed exec>:13, in <lambda>(x) File <timed exec>:9, in f(x) TypeError: 'bool' object does not support item assignment
错误原因
函数f(x)中参数x代表传入的当前行Series对象,但循环判断关键词时,执行了x = bool(re.search(element, x['text'], re.IGNORECASE)),直接把x从行对象覆盖成了布尔值。后续执行x['test'] = str(match)时,相当于对布尔类型对象做键赋值操作,而布尔类型不支持该操作,因此抛出对应错误。另外原代码将匹配结果转为字符串存储,和期望的布尔类型输出也不匹配。
修复方案
方案1:修正原apply逻辑
不要覆盖行对象变量x,将关键词匹配结果存入单独的临时变量,同时去掉多余的导入、提前赋值逻辑,直接存储布尔类型结果:
import pandas as pd import re d = { "customerId": [1, 3, 3], "text": ["Something with Cat", "That is a huge dog", "Hello agian"], } df = pd.DataFrame(data=d) my_list = ['cat', 'dog', 'mouse'] def f(x): match = False for element in my_list: # 匹配结果存入临时变量,不要覆盖行对象x is_match = bool(re.search(element, x['text'], re.IGNORECASE)) if is_match: match = True break x['test'] = match return x df = df.apply(lambda x: f(x), axis = 1)
方案2:向量化实现(性能更优)
逐行apply在数据量大时运行效率低,可直接用pandas字符串向量化方法实现,不需要手写循环:
# 将关键词拼接为正则或模式,忽略大小写一次性完成全列匹配 pattern = '|'.join(my_list) df['test'] = df['text'].str.contains(pattern, flags=re.IGNORECASE, na=False)
两种方案运行后都能得到符合预期的输出结果。
内容的提问来源于stack exchange,提问作者Test
相关产品推荐
相关产品推荐

