You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

词/句嵌入中discourse vector的含义、构建及相关疑问

Discourse Vectors: Definitions, Context in Arora’s Work, and Creation Methods

Great questions—let’s break this down clearly, especially tying back to Arora’s foundational papers since that’s where your confusion is rooted.

1. Core Definition of Discourse Vectors in Word/Sentence Embeddings

At its core, a discourse vector is a dense, low-dimensional numerical representation that captures the broader thematic or conversational scope of a text segment (sentence, paragraph, document, or longer discourse). Unlike standard word embeddings (which focus on individual word meanings) or basic sentence embeddings (which often just average word vectors), discourse vectors zero in on the unifying "topic or focus" that ties the text together—think of it as the "big picture" of what the text is about, beyond local word contexts.

2. Discourse Vectors in Arora’s Papers: Clarifying the Concept & Creation

Let’s unpack the two papers you mentioned, since the role and creation of discourse vectors shift slightly between them:

What Do They Represent? (Theme, Context, or Something Else?)

Arora’s line "discourse vector represents what is being talked about" is spot-on, but it’s helpful to narrow this down:

  • In both papers, it leans heavily toward thematic content—a "topic signature" for the text. It’s not just local context (like the words immediately surrounding a target term) but the global, overarching subject that unifies the discourse. For example, in a paragraph about renewable energy, the discourse vector would encode the core theme of renewable energy, rather than just the context around words like "solar" or "wind".
  • That said, it’s tied to context in the sense that it’s derived from statistical patterns of word usage across the discourse. So it’s a blend: a thematic representation grounded in how words co-occur globally within the text.

How Are They Created?

The method varies drastically between the two papers:

  • TACL 2016: A Latent Variable Model Approach to PMI-Based Word Embeddings
    Here, the discourse vector is a learned latent variable tied to each document:

    1. Start with a PMI (Pointwise Mutual Information) matrix, which captures how often words co-occur relative to their individual frequencies.
    2. Arora’s model decomposes this PMI matrix into two parts: word embeddings and discourse vectors.
    3. During training, the discourse vector is inferred alongside word embeddings by optimizing the model to fit the observed word co-occurrence data. It acts as a global bias that adjusts word embeddings to account for the document’s specific overarching topic.
  • ICLR 2017: A Simple but Tough-to-Beat Baseline for Sentence Embeddings
    In this paper, the discourse vector takes a simpler, fixed (corpus-specific) form:

    1. First, compute average word embeddings for every sentence in the training corpus.
    2. Calculate the first principal component (PC) of all these average sentence embeddings. This first PC is the discourse vector—it captures the most common global thematic variation across the corpus (think generic background noise or a default topic).
    3. To create a focused sentence embedding, subtract the projection of the sentence’s average embedding onto this fixed discourse vector. This removes redundant global thematic noise, leaving an embedding that highlights the sentence’s unique content.

Learned vs. Fixed?

It depends on the framework:

  • In the 2016 latent variable model: Learned per discourse segment (document). It’s a unique vector inferred for each document during training, tailored to its specific topic.
  • In the 2017 baseline: Fixed (corpus-specific). It’s computed once from the training corpus and reused for all sentences. It’s not learned per sentence, but it’s adapted to the corpus’s dominant themes.

内容的提问来源于stack exchange,提问作者Hyeseon Ko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:35:10