如何抑制NewsSentiment抛出的“Be aware, overflowing tokens are not returned”警告
解决NewsSentiment依赖Transformers时的截断策略警告
问题背景
使用NewsSentiment进行目标情感分析时,会触发Hugging Face Transformers的如下警告:
Be aware, overflowing tokens are not returned for the setting you have chosen, i.e. sequence pairs with the 'longest_first' truncation strategy. So the returned list will always be empty even if some tokens have been removed.
尝试过常规的warnings.filterwarnings配置但无效,原因是匹配规则不够精准,或者没有定位到抛出警告的具体模块。
有效抑制方案
方案1:精确匹配完整警告消息
之前的代码中消息文本被截断,导致匹配失败。使用完整的警告消息(或正则匹配前缀)来过滤:
import warnings # 完整匹配警告消息 warnings.filterwarnings( "ignore", message=r"Be aware, overflowing tokens are not returned for the setting you have chosen, i.e. sequence pairs with the 'longest_first' truncation strategy. So the returned list will always be empty even if some tokens have been removed." ) # 或者用正则前缀匹配(更灵活,避免标点/空格差异) warnings.filterwarnings( "ignore", message=r"Be aware, overflowing tokens are not returned for the setting you have chosen.*" )
方案2:精准定位Transformers的具体模块
该警告来自transformers.tokenization_utils_base模块,直接针对该模块的UserWarning过滤:
import warnings warnings.filterwarnings( "ignore", category=UserWarning, module=r"transformers\.tokenization_utils_base" )
方案3:从根源调整tokenizer配置(可选)
如果不想抑制警告,而是从逻辑上避免触发,可以获取NewsSentiment的tokenizer,修改截断策略为适合业务场景的选项:
from newsentiment import TargetSentiment ts = TargetSentiment() # 修改截断策略,避免触发警告 ts.tokenizer.truncation = "only_second" # 或"do_not_truncate"等其他策略
内容的提问来源于stack exchange,提问作者Steve
相关产品推荐
相关产品推荐

