You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网络分析中节点多变量关联对比与TS、PO、Art相关性计算咨询

解决方案:网络节点变量关联与相关性分析

一、统计拥有相同变量值对的节点数量

场景1:统计单个变量各取值对应的节点总数

如果需要统计每个变量(如TS)下,不同取值对应的节点数量,用pandas分组聚合即可实现:

import pandas as pd

# 读入CSV数据
df = pd.read_csv("your_data.csv")

# 统计TS变量各取值的节点数
ts_node_count = df.groupby("TS")["节点ID"].nunique()
print(ts_node_count)

# 同理批量处理PO、Art变量
for col in ["PO", "Art"]:
    count = df.groupby(col)["节点ID"].nunique()
    print(f"\n{col}变量各取值节点数:\n{count}")

场景2:统计两两节点间共享相同变量值的对数

如果需要统计任意两个节点在某变量上取值相同的总对数,可通过组合计数实现:

from itertools import combinations

# 按TS分组,计算每组内节点的两两组合数后求和
ts_same_pairs = df.groupby("TS")["节点ID"].apply(lambda x: len(list(combinations(x, 2)))).sum()
print(f"TS变量取值相同的节点对总数:{ts_same_pairs}")

# 同理处理PO、Art
po_same_pairs = df.groupby("PO")["节点ID"].apply(lambda x: len(list(combinations(x, 2)))).sum()
art_same_pairs = df.groupby("Art")["节点ID"].apply(lambda x: len(list(combinations(x, 2)))).sum()

二、计算TS、PO、Art三个变量的相关性

相关性计算需根据变量类型(数值型/分类型)选择对应方法,以下是具体实现:

情况1:变量为数值型(如TS是时间戳、PO是评分值)

用皮尔逊(线性相关)或斯皮尔曼(秩相关)系数衡量:

from scipy.stats import pearsonr, spearmanr

# 先清理缺失值
clean_df = df.dropna(subset=["TS", "PO", "Art"])

# 皮尔逊相关系数及显著性p值
ts_po_corr, ts_po_p = pearsonr(clean_df["TS"], clean_df["PO"])
ts_art_corr, ts_art_p = pearsonr(clean_df["TS"], clean_df["Art"])
po_art_corr, po_art_p = pearsonr(clean_df["PO"], clean_df["Art"])

print(f"TS与PO的皮尔逊相关系数:{ts_po_corr:.4f},p值:{ts_po_p:.4f}")
print(f"TS与Art的皮尔逊相关系数:{ts_art_corr:.4f},p值:{ts_art_p:.4f}")
print(f"PO与Art的皮尔逊相关系数:{po_art_corr:.4f},p值:{po_art_p:.4f}")

# 若变量非线性相关,改用斯皮尔曼秩相关
ts_po_spear, _ = spearmanr(clean_df["TS"], clean_df["PO"])
print(f"\nTS与PO的斯皮尔曼相关系数:{ts_po_spear:.4f}")

情况2:变量为分类型(如TS是类别标签、PO是分组ID)

用卡方检验、Cramér's V系数或互信息衡量关联强度:

from scipy.stats import chi2_contingency
import numpy as np
from sklearn.feature_selection import mutual_info_classif

clean_df = df.dropna(subset=["TS", "PO", "Art"])

# 卡方检验(以TS和PO为例)
contingency_table = pd.crosstab(clean_df["TS"], clean_df["PO"])
chi2, p, dof, expected = chi2_contingency(contingency_table)
print(f"TS与PO的卡方值:{chi2:.4f},p值:{p:.4f}")

# Cramér's V系数(修正卡方值,更适合分类变量关联)
n = contingency_table.sum().sum()
min_dim = min(contingency_table.shape) - 1
cramer_v = np.sqrt(chi2 / (n * min_dim))
print(f"TS与PO的Cramér's V系数:{cramer_v:.4f}")

# 互信息(衡量变量间依赖程度)
# 将分类变量转为数值编码
encoded_df = clean_df[["TS", "PO", "Art"]].apply(lambda x: pd.factorize(x)[0])
mi_ts_po = mutual_info_classif(encoded_df[["TS"]], encoded_df["PO"])[0]
print(f"TS与PO的互信息:{mi_ts_po:.4f}")

注意事项

  • 若CSV存在缺失值,必须先通过df.dropna()清理后再计算,避免结果失真
  • 若需结合网络节点的连接关系分析变量关联,可使用networkx库,比如统计相连节点间的变量取值相关性

内容的提问来源于stack exchange,提问作者user20564005

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 18:17:30