多语言词嵌入评估:Word Alignment与Dictionary Induction的差异咨询
Great question—this is a common point of confusion when diving into cross-lingual word embeddings, so let’s unpack it step by step.
These two tasks might sound similar, but their core goals and evaluation criteria are fundamentally distinct:
Word Alignment
This task focuses on local, sentence-level matching: given pairs of parallel sentences (e.g., a Chinese sentence and its English translation), you need to map individual words in one sentence to their direct semantic counterparts in the other. For example, in the pair "我 爱 猫" ↔ "I love cats", you’d align "我"→"I", "爱"→"love", "猫"→"cats".
- It relies heavily on contextual co-occurrence within a single sentence pair—the goal is to capture how words relate to each other in a specific linguistic context.
- Success is measured by how accurately you can match words that are directly paired in the same translated sentence.
Dictionary Induction
This task is about global, vocabulary-wide mapping: you aim to build a general-purpose dictionary that maps words from one language to their most semantically similar counterparts in another, regardless of whether they appear in the same parallel sentence. For example, creating an entry "猫"→"cat" that holds true across all contexts, not just a single sentence.
- It depends on global semantic similarity—the goal is to find words that have equivalent core meanings across the entire language, not just in a local context.
- Success is measured by how well your mappings align with human-curated dictionaries or standard reference datasets.
Now let’s connect this to the results you saw in the paper:
Bilingual Autoencoders Shine at Word Alignment
Bilingual autoencoders learn context-aware word representations. When encoding a word, they take into account the surrounding words in its sentence. This is perfect for word alignment, where you need to distinguish how a word’s meaning shifts based on its immediate context (e.g., differentiating "bank" as a financial institution vs. a river edge in different sentences). By modeling local context, these models can precisely match words that are paired in the same parallel sentence.
Inverted Index Excels at Dictionary Induction
Inverted indexes are built on global co-occurrence statistics: they count how often words from two languages appear together across all parallel corpus pairs, not just single sentences. This global view is ideal for dictionary induction, since it captures the most consistent, general-purpose semantic mappings between words. Unlike autoencoders, it doesn’t get distracted by context-specific nuances—instead, it focuses on the most frequent, stable word pairs that form the backbone of a cross-lingual dictionary.
The Tradeoff
- For word alignment, the inverted index’s global approach falls short: it can’t account for context-specific word relationships within individual sentences, leading to less precise local matches.
- For dictionary induction, the autoencoder’s context-aware representations become a liability: a word’s embedding shifts based on its sentence context, making it harder to find a single, consistent mapping that works across all uses of the word.
内容的提问来源于stack exchange,提问作者kpriya

