更新openpyxl后pandas调用any(1)突然报错求助
pandas过滤代码报错的解决方法
问题背景
更新openpyxl后,原本正常运行的pandas行过滤代码突然报错。
原代码及报错情况
原代码
data = {'Col1': ['Charges', 'Realized P&L', 'Other Credit & Debit', 'Some Other Value'], 'Col2': [100, 200, 300, 400], 'Col3': ['True', False, 'True', 'False']} df = pd.DataFrame(data) # 保留包含指定内容的行 filtered_df = df[df.isin(["Charges", "Realized P&L", "Other Credit & Debit"]).any(1)]
第一次报错信息
Traceback (most recent call last): File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.8/lib/python3.8/code.py", line 90, in runcode exec(code, self.locals) File "<input>", line 1, in <module> TypeError: any() takes 1 positional argument but 2 were given
修改后的代码(移除参数1)
filtered_df = df[df.isin(["Charges", "Realized P&L", "Other Credit & Debit"]).any()] # 移除了any()中的1
第二次报错信息
UserWarning: Boolean Series key will be reindexed to match DataFrame index. Traceback (most recent call last): File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.8/lib/python3.8/code.py", line 90, in runcode exec(code, self.locals) File "<input>", line 1, in <module> File "/Volumes/coding/venv/lib/python3.8/site-packages/pandas/core/frame.py", line 3751, in __getitem__ return self._getitem_bool_array(key) File "/Volumes/coding/venv/lib/python3.8/site-packages/pandas/core/frame.py", line 3804, in _getitem_bool_array key = check_bool_indexer(self.index, key) File "/Volumes/coding/venv/lib/python3.8/site-packages/pandas/core/indexing.py", line 2499, in check_bool_indexer raise IndexingError( pandas.errors.IndexingError: Unalignable boolean Series provided as indexer (index of the boolean Series and of the indexed object do not match).
问题原因
更新openpyxl时大概率连带升级了pandas版本。旧版pandas中any(1)是合法写法(用位置参数指定按行检查),但新版pandas对DataFrame.any()的参数做了规范,不再接受位置参数,必须显式指定axis=1。
移除参数1后,any()默认按列检查,返回的布尔Series长度等于列数,和原DataFrame的行数不匹配,因此出现索引对齐错误。
解决方案
将any(1)改为显式指定axis=1(或axis='columns',语义更清晰):
filtered_df = df[df.isin(["Charges", "Realized P&L", "Other Credit & Debit"]).any(axis=1)]
any(axis=1)会逐行检查布尔DataFrame中是否存在True值,生成一个和原DataFrame行数一致的布尔Series,用来过滤行即可正常运行。
内容的提问来源于stack exchange,提问作者Sid
相关产品推荐
相关产品推荐

