You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

向ZeroShotClassificationPipeline传入pandas列报错如何解决

报错根因

Hugging Face ZeroShotClassificationPipeline 只接受字符串、字符串列表、Hugging Face Dataset对象作为文本输入。直接传入pandas Series时,pipeline内部的输入校验逻辑会执行类似if input_data:的布尔判断,而pandas不允许直接对Series做整体布尔求值,就会抛出你看到的歧义错误。之前尝试的.apply()、.pipe()写法没生效,本质都是没有把输入转换成pipeline可识别的格式。

可直接运行的解决方案

中小数据集快速处理

直接把Series转成Python原生列表传入即可,写法最简单:

from transformers import pipeline
import pandas as pd

classifier = pipeline("zero-shot-classification", model="facebook/bart-large-mnli", device=0)
# 核心改动:用.tolist()把pandas列转成原生字符串列表
results = classifier(
    df['TEXT'].tolist(), 
    candidate_labels=label, 
    multi_label=True  # 旧版参数multi_class已更名,建议用新参数避免兼容问题
)

需要把结果写回原DataFrame

用.apply()的正确写法,逐行处理后直接存为新列:

def run_classify(single_text):
    pred = classifier(single_text, candidate_labels=label, multi_label=True)
    # 按需返回结果,示例是返回「标签:置信度」的字典,也可以只返回最高置信度的标签
    return dict(zip(pred['labels'], pred['scores']))

# 逐行推理,结果直接存入新列
df['cls_result'] = df['TEXT'].apply(run_classify)

大数据集高效推理

转成Hugging Face Dataset格式批量推理,速度比逐行apply快数倍:

from datasets import Dataset

# 从pandas生成Dataset对象
hf_dataset = Dataset.from_pandas(df[['TEXT']])
# 加batch_size参数开启批量推理,根据显存大小调整数值
results = classifier(
    hf_dataset['TEXT'], 
    candidate_labels=label, 
    multi_label=True,
    batch_size=16
)
注意事项
  • 禁止直接传入pandas Series/DataFrame对象给pipeline,必须做格式转换
  • 多标签分类场景优先用multi_label=True参数,旧的multi_class参数已经在新版本标记为弃用
  • 显存足够的话尽量调大batch_size,能大幅缩短推理时间

内容的提问来源于stack exchange,提问作者Lili.Y

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 17:42:14