You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于词向量在新数据集上构建情感分析分类器?

Using text2vec in R to Build Word Vectors for Feature Selection

Here's a structured workflow to build word vectors with R's text2vec package, designed to prepare meaningful features for a classification task. This follows the recommended practices for pre-classifier text vectorization from the text2vec project's official guidance.

Step 1: Load Dependencies and Sample Data

First, we'll load the required libraries and the built-in movie review dataset to work with:

# Loading packages and movie review data
require(text2vec)
require(data.table)
data("movie_review")
library(tidyverse)

Step 2: Clean and Structure the Dataset

Next, convert the review data into a data table (for efficient manipulation) and sort it by ID to maintain consistency across operations:

# Converting list of movie reviews to data table by reference
setDT(movie_review)

# Sorting the data table by ID
setkey(movie_review, id)

Step 3: Complete the Vectorization Workflow

Your original code snippet cuts off here, so here's how to finish the process to get word vectors and document-level features ready for classification:

3.1 Tokenize Text and Build a Clean Vocabulary

First, process the text into tokens and create a pruned vocabulary to remove rare, noisy terms:

# Tokenize the review text
tokens <- movie_review$review %>%
  tolower() %>%
  word_tokenizer()

# Create an iterator over tokens
it <- itoken(tokens, ids = movie_review$id)

# Build and prune vocabulary (keep words that appear at least 5 times)
vocab <- create_vocabulary(it)
vocab <- prune_vocabulary(vocab, term_count_min = 5)

3.2 Build Term-Co-Occurrence Matrix (TCM)

The TCM captures how often words appear near each other, which is critical for training semantically meaningful word vectors:

# Create vectorizer from the pruned vocabulary
vectorizer <- vocab_vectorizer(vocab)

# Build TCM with a skip-gram window of 5 (look 5 words before/after each term)
tcm <- create_tcm(it, vectorizer, skip_grams_window = 5)

3.3 Train GloVe Word Vectors

We use the GloVe algorithm to train word vectors that capture semantic relationships between terms:

# Initialize GloVe model
glove <- GlobalVectors$new(
  word_vectors_size = 50, # Size of each word vector
  vocabulary = vocab,
  x_max = 10 # Maximum frequency for weighting
)

# Train the model
wv_main <- glove$fit_transform(tcm, n_iter = 10, convergence_tol = 0.01)
wv_context <- glove$components
word_vectors <- wv_main + t(wv_context) # Combine main and context vectors

3.4 Create Document Vectors for Classification

Finally, convert each document into a single vector (average of its word vectors) to use as features for your classifier:

# Create Document-Term Matrix (DTM)
dtm <- create_dtm(it, vectorizer)

# Compute document vectors by multiplying DTM with word vectors
document_vectors <- dtm %*% word_vectors

These document_vectors can now be used as input features for your classification model (e.g., logistic regression, random forest). This approach leverages semantic information from the text, often leading to better model performance than traditional bag-of-words features.

内容的提问来源于stack exchange,提问作者Lucinho91

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:47:08