Mikolov 2014版Paragraph2Vec是否依赖段落顺序?含PV-DM/PV-DBOW及推特场景
1. Does Mikolov's 2014 Paragraph2Vec model assume sequential relationships within sentences/paragraphs?
Let's break this down by the two core variants of the model:
- PV-DM (Distributed Memory): Absolutely. This variant follows Word2Vec's CBOW approach, using a sliding window of words in their original order plus the paragraph vector to predict the central word. Word sequence is a critical part of how it ties paragraph context to word meaning.
- PV-DBOW (Distributed Bag of Words): Nope. This model ignores word order entirely. It uses the paragraph vector to predict random words sampled from the paragraph, focusing only on which words are present, not their sequence. Think of it as a bag-of-words take on paragraph embeddings.
2. Did the 2014 Paragraph2Vec paper assume valid paragraph order for both PV-DM and PV-DBOW? Will embeddings work for unordered Twitter streams?
First off: the paper does not assume any meaningful sequential relationship between different paragraphs for either model. When training, you feed paragraphs in some order, but the model never uses that inter-paragraph sequence to learn representations. Each paragraph vector is linked only to its own content, not the tweets (or paragraphs) before or after it.
For unordered Twitter streams (each tweet = one paragraph):
- PV-DM will still produce useful embeddings. Even though tweets are short, it captures local word order patterns within each individual tweet, which can add subtle semantic context.
- PV-DBOW is actually ideal here. Since it doesn't care about word order, it handles the short, informal nature of tweets perfectly, focusing on the overall meaning of each tweet rather than syntactic structure.
Bottom line: Training on an unordered set of tweets will absolutely yield valid, semantically meaningful paragraph vectors. The model's whole point is to map each piece of text (tweet) to a vector that reflects its content—regardless of how you shuffle the input paragraphs during training.
内容的提问来源于stack exchange,提问作者Franklin Dong

