使用Hugging Face训练的BERT模型批量预测推特文本情感
推特文本BERT批量情感预测实现方案
背景说明
我正在开展推特文本情感分析工作,需将情感划分为正向、负向、中性三类。已通过Hugging Face框架训练好BERT模型,需要对数据框中无标签的推特文本列进行批量情感预测,最初没有明确的实现思路。
我参考的教程仅提供了单条原始文本的预测示例,可正常运行的单条预测代码如下:
review_text = "I love completing my todos! Best app ever!!!" encoded_review = tokenizer.encode_plus( review_text, max_length=MAX_LEN, add_special_tokens=True, return_token_type_ids=False, pad_to_max_length=True, return_attention_mask=True, return_tensors='pt', ) input_ids = encoded_review['input_ids'].to(device) attention_mask = encoded_review['attention_mask'].to(device) output = model(input_ids, attention_mask) _, prediction = torch.max(output, dim=1) print(f'Review text: {review_text}') print(f'Sentiment : {class_names[prediction]}') # 输出示例 # Review text: I love completing my todos! Best app ever!!! # Sentiment : positive
批量预测可运行解决方案
经测试可正常使用的批量预测代码如下,核心逻辑是将单条预测流程封装为自定义函数,再通过pandas的apply方法批量作用于目标文本列:
def predictionPipeline(text): encoded_review = tokenizer.encode_plus( text, max_length=MAX_LEN, add_special_tokens=True, return_token_type_ids=False, pad_to_max_length=True, return_attention_mask=True, return_tensors='pt', ) input_ids = encoded_review['input_ids'].to(device) attention_mask = encoded_review['attention_mask'].to(device) output = model(input_ids, attention_mask) _, prediction = torch.max(output, dim=1) return(class_names[prediction]) df2['prediction']=df2['cleaned_tweet'].apply(predictionPipeline)
内容的提问来源于stack exchange,提问作者CLopez138
相关产品推荐
相关产品推荐

