如何基于词向量在新数据集上构建情感分析分类器?
Here's a structured workflow to build word vectors with R's text2vec package, designed to prepare meaningful features for a classification task. This follows the recommended practices for pre-classifier text vectorization from the text2vec project's official guidance.
Step 1: Load Dependencies and Sample Data
First, we'll load the required libraries and the built-in movie review dataset to work with:
# Loading packages and movie review data require(text2vec) require(data.table) data("movie_review") library(tidyverse)
Step 2: Clean and Structure the Dataset
Next, convert the review data into a data table (for efficient manipulation) and sort it by ID to maintain consistency across operations:
# Converting list of movie reviews to data table by reference setDT(movie_review) # Sorting the data table by ID setkey(movie_review, id)
Step 3: Complete the Vectorization Workflow
Your original code snippet cuts off here, so here's how to finish the process to get word vectors and document-level features ready for classification:
3.1 Tokenize Text and Build a Clean Vocabulary
First, process the text into tokens and create a pruned vocabulary to remove rare, noisy terms:
# Tokenize the review text tokens <- movie_review$review %>% tolower() %>% word_tokenizer() # Create an iterator over tokens it <- itoken(tokens, ids = movie_review$id) # Build and prune vocabulary (keep words that appear at least 5 times) vocab <- create_vocabulary(it) vocab <- prune_vocabulary(vocab, term_count_min = 5)
3.2 Build Term-Co-Occurrence Matrix (TCM)
The TCM captures how often words appear near each other, which is critical for training semantically meaningful word vectors:
# Create vectorizer from the pruned vocabulary vectorizer <- vocab_vectorizer(vocab) # Build TCM with a skip-gram window of 5 (look 5 words before/after each term) tcm <- create_tcm(it, vectorizer, skip_grams_window = 5)
3.3 Train GloVe Word Vectors
We use the GloVe algorithm to train word vectors that capture semantic relationships between terms:
# Initialize GloVe model glove <- GlobalVectors$new( word_vectors_size = 50, # Size of each word vector vocabulary = vocab, x_max = 10 # Maximum frequency for weighting ) # Train the model wv_main <- glove$fit_transform(tcm, n_iter = 10, convergence_tol = 0.01) wv_context <- glove$components word_vectors <- wv_main + t(wv_context) # Combine main and context vectors
3.4 Create Document Vectors for Classification
Finally, convert each document into a single vector (average of its word vectors) to use as features for your classifier:
# Create Document-Term Matrix (DTM) dtm <- create_dtm(it, vectorizer) # Compute document vectors by multiplying DTM with word vectors document_vectors <- dtm %*% word_vectors
These document_vectors can now be used as input features for your classification model (e.g., logistic regression, random forest). This approach leverages semantic information from the text, often leading to better model performance than traditional bag-of-words features.
内容的提问来源于stack exchange,提问作者Lucinho91

