You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用MathNet.Numerics计算Cramer's V矩阵?

用MathNet.Numerics计算Cramer's V矩阵分析类别共现关联

步骤1:数据预处理

先把时间序列中的类别数据转换成适合计算的格式,提取唯一类别并统计每对类别的共现情况:

示例数据整理

给定的6个时间点类别集合:

  • 6月14日上午:{香蕉, 番茄}
  • 6月14日下午:{牛奶, 豆子, 香蕉}
  • 6月15日上午:{苹果, 肉类, 番茄}
  • 6月15日下午:{鸡肉, 香蕉, 咖啡}
  • 6月16日上午:{牛奶, 豆子, 鸡肉}
  • 6月16日下午:{番茄, 橙子, 咖啡}

唯一类别列表:["香蕉", "番茄", "牛奶", "豆子", "苹果", "肉类", "鸡肉", "咖啡", "橙子"]

构建两两类别列联表

对每一对类别(A,B),构建2×2列联表,统计4种情况的样本数:

  • a:同时出现A和B的样本数
  • b:出现B但不出现A的样本数
  • c:出现A但不出现B的样本数
  • d:既不出现A也不出现B的样本数

比如香蕉和番茄的列联表:

出现香蕉不出现香蕉
出现番茄12
不出现番茄21

步骤2:安装MathNet.Numerics

通过NuGet安装库:

Install-Package MathNet.Numerics

或用.NET CLI:

dotnet add package MathNet.Numerics

步骤3:计算Cramer's V矩阵

以下是C#实现代码,遍历所有类别对计算每对的关联度:

using System;
using System.Collections.Generic;
using System.Linq;
using MathNet.Numerics.Statistics;

class CramersVCalculator
{
    static void Main()
    {
        // 定义样本数据:每个元素是一个时间点的类别集合
        var samples = new List<HashSet<string>>
        {
            new HashSet<string> {"香蕉", "番茄"},
            new HashSet<string> {"牛奶", "豆子", "香蕉"},
            new HashSet<string> {"苹果", "肉类", "番茄"},
            new HashSet<string> {"鸡肉", "香蕉", "咖啡"},
            new HashSet<string> {"牛奶", "豆子", "鸡肉"},
            new HashSet<string> {"番茄", "橙子", "咖啡"}
        };

        // 获取所有唯一类别并排序
        var categories = samples.SelectMany(s => s).Distinct().OrderBy(c => c).ToList();
        int categoryCount = categories.Count;

        // 初始化Cramer's V矩阵
        double[,] cramersVMatrix = new double[categoryCount, categoryCount];

        // 遍历所有类别对计算V值
        for (int i = 0; i < categoryCount; i++)
        {
            string catA = categories[i];
            // 对角线设为1(自身关联度为1)
            cramersVMatrix[i, i] = 1.0;

            for (int j = i + 1; j < categoryCount; j++)
            {
                string catB = categories[j];

                // 构建2×2列联表
                int a = samples.Count(s => s.Contains(catA) && s.Contains(catB));
                int b = samples.Count(s => !s.Contains(catA) && s.Contains(catB));
                int c = samples.Count(s => s.Contains(catA) && !s.Contains(catB));
                int d = samples.Count(s => !s.Contains(catA) && !s.Contains(catB));
                int n = samples.Count;

                // 计算卡方值(避免除零)
                double chiSquared = 0;
                double denominator = (a + b) * (c + d) * (a + c) * (b + d);
                if (denominator != 0)
                {
                    chiSquared = n * Math.Pow(a * d - b * c, 2) / denominator;
                }

                // 2×2表中Cramer's V简化为√(χ²/n)
                double cramersV = Math.Sqrt(chiSquared / n);

                // 矩阵对称,填充对称位置
                cramersVMatrix[i, j] = cramersV;
                cramersVMatrix[j, i] = cramersV;
            }
        }

        // 输出结果矩阵
        Console.WriteLine("Cramer's V矩阵:");
        Console.Write("\t");
        foreach (var cat in categories)
        {
            Console.Write($"{cat}\t");
        }
        Console.WriteLine();

        for (int i = 0; i < categoryCount; i++)
        {
            Console.Write($"{categories[i]}\t");
            for (int j = 0; j < categoryCount; j++)
            {
                Console.Write($"{cramersVMatrix[i, j]:F4}\t");
            }
            Console.WriteLine();
        }
    }
}

代码说明

  • 用HashSet存储样本类别,方便快速判断类别是否存在
  • 对每对类别构建2×2列联表,统计四种共现情况
  • 卡方值计算遵循标准公式,处理分母为零的边界情况
  • 矩阵为对称矩阵,对角线值固定为1(类别自身完全关联)

结果解读

Cramer's V值范围在0到1之间:

  • 0表示两个类别完全无关联
  • 1表示两个类别完全关联(总是同时出现或同时不出现)

比如示例中,牛奶和豆子的V值约为0.8165,说明二者共现概率很高;香蕉和橙子的V值为0,说明从未同时出现。

内容的提问来源于stack exchange,提问作者Farzad M.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 16:04:50