You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python的多变量分析:双变量响应分布度量需求

分析建议:描述2D直方图中的响应“团块”特征

你的场景里,线性拟合或普通相关性(比如Pearson)确实不适合——这类指标衡量的是全局线性趋势,而你要描述的是响应变量在(x,y)空间的局部高响应聚集区域,全局指标会被非团块区域的信息稀释,完全捕捉不到团块的形状、密度、强度这些核心特征。

下面是几个更贴合需求的分析思路,附Python实现示例:

1. 基于直方图网格的连通区域分析(适合已生成的2D直方图)

如果你的输入是已经分好网格的2D直方图数据,核心思路是先筛选出高响应的网格,再标记连通的团块,然后计算团块的关键特征:

  • 峰值/平均响应强度
  • 团块覆盖的网格面积
  • 紧凑度(衡量团块的规整程度)
  • 团块的坐标范围
import numpy as np
from scipy.ndimage import label, find_objects

# 假设hist是你的2D直方图数据,shape=(nx, ny),值代表响应强度(浅色对应高值)
# 第一步:设定阈值筛选高响应区域(可调整百分位数,比如取前20%的高响应)
response_threshold = np.percentile(hist, 80)
high_response_mask = hist > response_threshold

# 第二步:标记连通的团块(8邻域或4邻域可通过structure参数调整)
labeled_clusters, cluster_count = label(high_response_mask)

# 第三步:遍历每个团块计算特征
for cluster_id in range(1, cluster_count + 1):
    # 获取当前团块的边界切片
    cluster_slice = find_objects(labeled_clusters == cluster_id)[0]
    cluster_data = hist[cluster_slice]
    
    # 计算特征
    peak_intensity = np.max(cluster_data)
    mean_intensity = np.mean(cluster_data)
    grid_area = np.sum(labeled_clusters == cluster_id)
    # 紧凑度:团块实际覆盖网格数 / 外接矩形网格数
    rect_grid_count = (cluster_slice[0].stop - cluster_slice[0].start) * (cluster_slice[1].stop - cluster_slice[1].start)
    compactness = grid_area / rect_grid_count
    
    print(f"团块{cluster_id}: 峰值强度={peak_intensity:.2f}, 平均强度={mean_intensity:.2f}, 覆盖网格数={grid_area}, 紧凑度={compactness:.2f}")

2. 基于原始样本的密度聚类分析(适合未分网格的原始数据)

如果你有(x,y,响应值)的原始样本点,可以直接筛选高响应样本,用密度聚类(比如DBSCAN)识别团块,然后提取聚类特征:

  • 团块的样本数量
  • 团块的中心坐标
  • 团块内的平均/峰值响应
import numpy as np
from sklearn.cluster import DBSCAN
from sklearn.preprocessing import StandardScaler
import pandas as pd

# 假设data是包含x、y、response列的DataFrame
data = pd.DataFrame(...)  # 替换为你的原始数据
# 筛选高响应样本(同样可调整百分位数)
high_response_samples = data[data['response'] > np.percentile(data['response'], 80)]

# 标准化x、y变量(消除量纲影响)
scaled_coords = StandardScaler().fit_transform(high_response_samples[['x', 'y']])

# DBSCAN聚类:eps是邻域半径,min_samples是形成簇的最小样本数,可根据数据调整
dbscan = DBSCAN(eps=0.5, min_samples=5)
cluster_labels = dbscan.fit_predict(scaled_coords)

# 遍历每个有效簇(排除噪声点label=-1)
unique_labels = set(cluster_labels)
for label in unique_labels:
    if label == -1:
        continue
    cluster = high_response_samples[cluster_labels == label]
    print(f"簇{label}: 样本数={len(cluster)}, 中心坐标=({cluster['x'].mean():.2f}, {cluster['y'].mean():.2f}), 平均响应={cluster['response'].mean():.2f}")

3. 空间自相关分析(衡量团块的显著性)

如果需要验证团块是否是显著的聚集(而非随机分布),可以用局部Moran's I统计量,识别显著的高-高聚集区域:

import libpysal
from esda.moran import Moran_Local
import numpy as np

# 假设hist是2D直方图数据
hist_flat = hist.flatten()
# 构建网格空间权重矩阵(邻接关系)
spatial_weights = libpysal.weights.lat2W(hist.shape[0], hist.shape[1])

# 计算局部Moran's I
local_moran = Moran_Local(hist_flat, spatial_weights)
# 提取p值<0.05的显著高-高聚集区域(q=1代表高响应聚集)
significant_clusters = (local_moran.p_sim < 0.05) & (local_moran.q == 1)
# 转回2D矩阵查看分布
significant_clusters_2d = significant_clusters.reshape(hist.shape)

总结

  • 线性拟合/R²这类全局指标完全不匹配你的需求,它们无法捕捉局部聚集的特征;
  • 优先根据数据类型选择方法:已有直方图用连通区域分析,原始样本用密度聚类;
  • 若需要验证团块的统计显著性,可搭配空间自相关分析。

内容的提问来源于stack exchange,提问作者WolfiG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.11 23:43:12