使用warnings.filterwarnings()过滤PyTables的PerformanceWarning失效
如何抑制PyTables的PerformanceWarning?
常规的Python忽略特定警告方案(比如按类别过滤、正则匹配警告消息)在抑制PyTables的PerformanceWarning时完全无效,以下是具体场景及尝试过的方法:
最小可复现示例(MWE)
import pandas as pd import warnings from tables import NaturalNameWarning, PerformanceWarning data = { 'a' : 1, 'b' : 'two' } df = pd.DataFrame.from_dict(data, orient = 'index') # 混合类型会触发PerformanceWarning dest = pd.HDFStore('warnings.h5', 'w') # dest.put('data', df) # 混合类型会产生PerformanceWarning # dest.put('data 1', df) # 'data 1'中的空格会同时触发NaturalNameWarning和PerformanceWarning warnings.filterwarnings('ignore', category = NaturalNameWarning) # NaturalNameWarning可以被忽略 warnings.filterwarnings('ignore', category = PerformanceWarning) # 无效果 warnings.filterwarnings('ignore', message='.*PyTables will pickle') # 无效果 # warnings.filterwarnings('ignore') # 会屏蔽所有警告,不符合需求 dest.put('data 2', df) # 仍会弹出PerformanceWarning dest.close()
尝试过的其他无效方法
- 使用上下文管理器:
with warnings.catch_warnings(): warnings.filterwarnings("ignore", category=PerformanceWarning) # 无效果 warnings.filterwarnings('ignore', message='.*PyTables') # 无效果 dest.put('data 6', df)
- 改用
warnings.simplefilter()替代warnings.filterwarnings(),同样无效。
警告信息详情
PerformanceWarning
test.py:21: PerformanceWarning: your performance may suffer as PyTables will pickle object types that it cannot map directly to c-types [inferred_type->mixed-integer,key->block0_values] [items->Int64Index([0], dtype='int64')] dest.put('data 2', df) # PerformanceWarning
NaturalNameWarning(可正常被忽略)
/home/user/.local/lib/python3.8/site-packages/tables/path.py:137: NaturalNameWarning: object name is not a valid Python identifier: 'data 2'; it does not match the pattern ``^[a-zA-Z_][a-zA-Z0-9_]*$``; you will not be able to use natural naming to access this object; using ``getattr()`` will still work, though check_attribute_name(name)
当前环境:tables 3.7.0 / Python 3.8.10
解决思路
1. 根源规避:指定HDF存储格式为table
Pandas的HDFStore.put默认用fixed格式,混合类型会触发PyTables的警告。改用format='table'可以直接绕过这个问题,同时还支持更灵活的后续查询操作,这是最推荐的方案:
dest.put('data 2', df, format='table') # 不会触发PerformanceWarning
2. 前置全局过滤警告
在脚本最开头就设置警告过滤器,确保在导入pandas和PyTables之前生效:
import warnings # 先设置过滤器再导入相关库 from tables import PerformanceWarning, NaturalNameWarning warnings.simplefilter('ignore', PerformanceWarning) warnings.simplefilter('ignore', NaturalNameWarning) # 之后再导入pandas并执行逻辑 import pandas as pd data = {'a':1, 'b':'two'} df = pd.DataFrame.from_dict(data, orient='index') dest = pd.HDFStore('warnings.h5', 'w') dest.put('data 2', df) dest.close()
3. 上下文管理器内精准过滤
在上下文管理器内重新导入警告类并设置过滤器,确保作用域生效:
import pandas as pd import warnings data = {'a':1, 'b':'two'} df = pd.DataFrame.from_dict(data, orient='index') dest = pd.HDFStore('warnings.h5', 'w') with warnings.catch_warnings(): from tables import PerformanceWarning, NaturalNameWarning warnings.filterwarnings("ignore", category=NaturalNameWarning) warnings.filterwarnings("ignore", category=PerformanceWarning) dest.put('data 2', df) dest.close()
4. 临时修改PyTables源码(不推荐)
找到PyTables中抛出PerformanceWarning的源码文件(通常在tables/array.py或相关IO模块),注释掉对应的warnings.warn()代码。但该方法依赖当前版本,升级PyTables后会失效。
内容的提问来源于stack exchange,提问作者the.real.gruycho
相关产品推荐
相关产品推荐

