You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Numpy(非Pandas)实现基于窗口的词语共现矩阵?

用Numpy实现词语共现矩阵(窗口式)

核心步骤拆解

1. 构建词-索引映射

Numpy矩阵靠索引定位元素,所以首先要把所有唯一词汇映射到整数索引:

  • 先把所有句子分词,收集所有不重复的词汇
  • 用字典建立词汇到索引的对应关系,同时记录词汇总数

2. 初始化共现矩阵

创建一个大小为词汇数×词汇数的Numpy数组,初始值全为0,对角线自然为0(后续只有当词在自身窗口重复时才会修改)。

3. 遍历句子与窗口,更新矩阵

对每个句子的每个词,确定其窗口的有效范围(处理边界),然后遍历窗口内的所有上下文词,在矩阵对应位置累加计数。


完整代码实现

import numpy as np

# 输入句子
sentences = ["All that glitters is not gold", "All that I want is You"]
window_size = 4

# 1. 分词并收集唯一词汇
tokenized_sentences = [sent.split() for sent in sentences]
all_words = []
for tokens in tokenized_sentences:
    for word in tokens:
        if word not in all_words:
            all_words.append(word)

# 构建词-索引映射
word_to_idx = {word: idx for idx, word in enumerate(all_words)}
vocab_size = len(all_words)

# 2. 初始化共现矩阵
cooccurrence_matrix = np.zeros((vocab_size, vocab_size), dtype=np.int32)

# 3. 遍历更新矩阵
for tokens in tokenized_sentences:
    sent_length = len(tokens)
    for target_pos, target_word in enumerate(tokens):
        target_idx = word_to_idx[target_word]
        
        # 处理左侧窗口:从max(0, 当前位置-窗口大小)到当前位置前一位
        left_start = max(0, target_pos - window_size)
        for context_pos in range(left_start, target_pos):
            context_word = tokens[context_pos]
            context_idx = word_to_idx[context_word]
            cooccurrence_matrix[target_idx][context_idx] += 1
        
        # 处理右侧窗口:从当前位置后一位到min(句末, 当前位置+窗口大小)
        right_end = min(sent_length - 1, target_pos + window_size)
        for context_pos in range(target_pos + 1, right_end + 1):
            context_word = tokens[context_pos]
            context_idx = word_to_idx[context_word]
            cooccurrence_matrix[target_idx][context_idx] += 1

# 打印结果(可选)
print("词汇列表:", all_words)
print("共现矩阵:")
print(cooccurrence_matrix)

关键细节说明

  1. 边界处理:

    • 左侧窗口用max(0, target_pos - window_size)确保不会超出句子开头
    • 右侧窗口用min(sent_length - 1, target_pos + window_size)确保不会超出句子结尾
      比如示例中第一句的最后一个词"gold"(位置5),右侧没有词,所以只处理左侧窗口(位置5-4=1到4),即与"glitters"、"is"、"not"共现。
  2. 词-索引映射:
    用字典word_to_idx快速查找词汇对应的矩阵行列索引,这是Numpy矩阵操作的核心前提——把字符串词汇转化为可计算的整数索引。

  3. 对角线处理:
    因为循环中上下文位置始终不等于目标词位置(context_pos不会等于target_pos),所以对角线默认保持0。如果句子中出现重复词在窗口内(比如["All All that"]),则会触发对角线位置的计数累加。

  4. 共现的双向性:
    上述代码实现的是单向计数(比如"All"与"that"共现,只在"All"行、"that"列加1)。如果需要对称共现矩阵(即"that"行、"All"列也加1),只需在更新时同时执行cooccurrence_matrix[context_idx][target_idx] += 1即可。

内容的提问来源于stack exchange,提问作者Tababy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 05:01:02