You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Hugging Face进行ABSA遇阻:求方面与情感提取逻辑

解决方案

方式一:使用Hugging Face Token分类Pipeline(推荐新手)

该模型属于token级别的分类任务,专门用于识别句子中的方面词并对应情感,用token-classification pipeline可以自动处理分词、标签聚合,代码更简洁:

from transformers import pipeline

model_name = "yangheng/deberta-v3-base-absa-v1.1"
# 启用聚合策略,自动合并连续的方面词token
absa_pipeline = pipeline("token-classification", model=model_name, tokenizer=model_name, aggregation_strategy="simple")

text = "The food was great but the service was terrible."
results = absa_pipeline(text)

print(f"text - {text}\n")
aspect_sentiment_map = {}
current_aspect = []
current_sentiment = None

for entity in results:
    if entity['entity_group'] == 'Aspect':
        # 去除分词生成的子词标记(如##)
        clean_word = entity['word'].replace(' ##', '')
        current_aspect.append(clean_word)
    elif entity['entity_group'] in ['Positive', 'Negative']:
        current_sentiment = entity['entity_group']
        aspect = ' '.join(current_aspect).strip()
        aspect_sentiment_map[aspect] = current_sentiment.lower()
        current_aspect = []

# 输出最终结果
for aspect, sentiment in aspect_sentiment_map.items():
    print(f"aspect - {aspect}, sentiment {sentiment}")

方式二:手动解析模型输出(理解底层逻辑)

如果想手动处理张量,需要按照模型的标签逻辑解析token级预测结果:

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model_name = "yangheng/deberta-v3-base-absa-v1.1"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
text = "The food was great but the service was terrible."

# 输入处理,保留offset映射用于对齐原文本与token
inputs = tokenizer(text, return_tensors="pt", return_offsets_mapping=True)
offset_mapping = inputs.pop('offset_mapping').squeeze().tolist()
outputs = model(**inputs)

# 获取每个token的预测标签
predictions = torch.argmax(outputs.logits, dim=2).squeeze().tolist()
label_map = model.config.id2label

aspect_sentiment = {}
current_aspect = []

for idx, pred_id in enumerate(predictions):
    # 跳过CLS和SEP特殊token
    if idx == 0 or idx == len(predictions)-1:
        continue
    
    label = label_map[pred_id]
    start, end = offset_mapping[idx]
    token_text = text[start:end]

    if label == 'B-Aspect':
        current_aspect = [token_text]
    elif label == 'I-Aspect':
        current_aspect.append(token_text)
    elif label in ['Positive', 'Negative']:
        aspect = ''.join(current_aspect).strip()
        aspect_sentiment[aspect] = label.lower()
        current_aspect = []

# 输出结果
print(f"text - {text}\n")
for aspect, sentiment in aspect_sentiment.items():
    print(f"aspect - {aspect}, sentiment {sentiment}")

关键修正点

  • 该模型的核心逻辑是先识别方面词的范围(B-Aspect/I-Aspect标签),再匹配对应的情感标签(Positive/Negative),而非直接给每个token分配情感。
  • 之前的错误在于将模型输出与单个token直接绑定,忽略了模型的token级分类任务设定。

内容的提问来源于stack exchange,提问作者Dexter1611

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 05:45:31