使用Hugging Face进行ABSA遇阻:求方面与情感提取逻辑
解决方案
方式一:使用Hugging Face Token分类Pipeline(推荐新手)
该模型属于token级别的分类任务,专门用于识别句子中的方面词并对应情感,用token-classification pipeline可以自动处理分词、标签聚合,代码更简洁:
from transformers import pipeline model_name = "yangheng/deberta-v3-base-absa-v1.1" # 启用聚合策略,自动合并连续的方面词token absa_pipeline = pipeline("token-classification", model=model_name, tokenizer=model_name, aggregation_strategy="simple") text = "The food was great but the service was terrible." results = absa_pipeline(text) print(f"text - {text}\n") aspect_sentiment_map = {} current_aspect = [] current_sentiment = None for entity in results: if entity['entity_group'] == 'Aspect': # 去除分词生成的子词标记(如##) clean_word = entity['word'].replace(' ##', '') current_aspect.append(clean_word) elif entity['entity_group'] in ['Positive', 'Negative']: current_sentiment = entity['entity_group'] aspect = ' '.join(current_aspect).strip() aspect_sentiment_map[aspect] = current_sentiment.lower() current_aspect = [] # 输出最终结果 for aspect, sentiment in aspect_sentiment_map.items(): print(f"aspect - {aspect}, sentiment {sentiment}")
方式二:手动解析模型输出(理解底层逻辑)
如果想手动处理张量,需要按照模型的标签逻辑解析token级预测结果:
from transformers import AutoTokenizer, AutoModelForSequenceClassification import torch model_name = "yangheng/deberta-v3-base-absa-v1.1" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForSequenceClassification.from_pretrained(model_name) text = "The food was great but the service was terrible." # 输入处理,保留offset映射用于对齐原文本与token inputs = tokenizer(text, return_tensors="pt", return_offsets_mapping=True) offset_mapping = inputs.pop('offset_mapping').squeeze().tolist() outputs = model(**inputs) # 获取每个token的预测标签 predictions = torch.argmax(outputs.logits, dim=2).squeeze().tolist() label_map = model.config.id2label aspect_sentiment = {} current_aspect = [] for idx, pred_id in enumerate(predictions): # 跳过CLS和SEP特殊token if idx == 0 or idx == len(predictions)-1: continue label = label_map[pred_id] start, end = offset_mapping[idx] token_text = text[start:end] if label == 'B-Aspect': current_aspect = [token_text] elif label == 'I-Aspect': current_aspect.append(token_text) elif label in ['Positive', 'Negative']: aspect = ''.join(current_aspect).strip() aspect_sentiment[aspect] = label.lower() current_aspect = [] # 输出结果 print(f"text - {text}\n") for aspect, sentiment in aspect_sentiment.items(): print(f"aspect - {aspect}, sentiment {sentiment}")
关键修正点
- 该模型的核心逻辑是先识别方面词的范围(B-Aspect/I-Aspect标签),再匹配对应的情感标签(Positive/Negative),而非直接给每个token分配情感。
- 之前的错误在于将模型输出与单个token直接绑定,忽略了模型的token级分类任务设定。
内容的提问来源于stack exchange,提问作者Dexter1611
相关产品推荐
相关产品推荐

