如何用Keras的Embedding层实现词转向量及生成自定义数据集词嵌入
Hey there! Let's break down your two questions clearly—how to generate word embeddings for your own dataset, and how to use Keras' Embedding layer to turn words into vectors.
Here's a step-by-step workflow to create word embeddings tailored to your data:
Step 1: Text Preprocessing
- Clean your text first: Remove punctuation, special characters, normalize case (e.g., all lowercase), and optionally filter stopwords (like "the", "in") depending on your task.
- Tokenize the text: Split sentences into individual words (for English this is often space-based; for Chinese you’ll need dedicated tokenizers).
- Build a vocabulary: Collect all unique words from your corpus, assign each a unique integer ID (e.g., "the" → 1, "in" → 2). You can also filter out rare words to reduce noise.
Step 2: Choose a Training Approach
You have two main options:- Use Pre-trained Embeddings: Leverage existing models trained on large corpora (like Word2Vec, GloVe, or BERT embeddings). This is great if you have a small dataset, as pre-trained vectors capture general language patterns.
- Train Custom Embeddings: If your data is domain-specific (e.g., medical, legal text), train your own embeddings:
- Use architectures like Skip-gram or CBOW (from Word2Vec) to learn vectors by predicting context words.
- Or train embeddings alongside your downstream task (e.g., text classification, NER) using deep learning models like LSTMs or Transformers—this lets the vectors adapt to your specific task needs.
Step 3: Load or Train the Embeddings
- For pre-trained embeddings: Download the embedding file (in the format you provided—one word per line followed by its vector), then map the words to your vocabulary. Words not in the pre-trained set can be assigned random vectors or the average of all vectors.
- For custom training: Build your model, feed in your tokenized text, and train until the embeddings stabilize. Save the final vectors in a format like your example (word + space-separated values).
Step 4: Validate and Refine
- Check embedding quality by calculating word similarity (e.g., cosine similarity). For example,
king - man + womanshould be close toqueenfor well-trained embeddings. - Adjust preprocessing steps, model parameters, or training data based on your task’s performance.
- Check embedding quality by calculating word similarity (e.g., cosine similarity). For example,
Keras' Embedding layer simplifies turning integer-encoded words into vectors. You can use it with pre-trained embeddings or train from scratch.
1. Prepare Data: Convert Words to Integer IDs
First, turn your text into sequences of integers using Keras' Tokenizer:
from tensorflow.keras.preprocessing.text import Tokenizer from tensorflow.keras.preprocessing.sequence import pad_sequences # Example text data texts = ["the cat in the hat", "to be or not to be"] # Initialize tokenizer (keep top 1000 most frequent words) tokenizer = Tokenizer(num_words=1000) tokenizer.fit_on_texts(texts) # Convert text to integer sequences sequences = tokenizer.texts_to_sequences(texts) # Pad sequences to a fixed length (required for model input) padded_sequences = pad_sequences(sequences, maxlen=10) # Get the word-to-ID mapping word_index = tokenizer.word_index
2. Option 1: Load Pre-trained Embeddings into the Layer
If you have a pre-trained embedding file (like the one you shared), load it into an embedding matrix and pass it to the Embedding layer:
import numpy as np # Define embedding dimensions (matches your file's vector size) embedding_dim = 5 # Create a matrix to hold embeddings (index 0 is reserved for padding) embedding_matrix = np.zeros((len(word_index) + 1, embedding_dim)) # Load the pre-trained file with open("your_embedding_file.txt", encoding="utf-8") as f: for line in f: word, vec_str = line.split(maxsplit=1) vec = np.fromstring(vec_str, sep=" ") if word in word_index: idx = word_index[word] embedding_matrix[idx] = vec # Build the Embedding layer (set trainable=False to keep pre-trained weights fixed) from tensorflow.keras.layers import Embedding embedding_layer = Embedding( input_dim=len(word_index) + 1, output_dim=embedding_dim, weights=[embedding_matrix], input_length=10, # Matches padded_sequences' maxlen trainable=False )
3. Option 2: Train Embeddings from Scratch
If you don’t have pre-trained embeddings, train them directly with your task:
embedding_layer = Embedding( input_dim=len(word_index) + 1, output_dim=128, # Choose your desired vector dimension input_length=10 )
Add this layer to your model (e.g., paired with an LSTM or Dense layer), and the embeddings will be optimized as you train your downstream task.
4. Convert Integer Sequences to Vectors
Once your Embedding layer is set up, use it to generate vectors:
import tensorflow as tf # Convert padded sequences to a tensor input_tensor = tf.convert_to_tensor(padded_sequences) # Get the embedded vectors embedded_vectors = embedding_layer(input_tensor) # Output shape: (number of samples, sequence length, embedding dimension) print(embedded_vectors.shape)
内容的提问来源于stack exchange,提问作者Jose Kj

