生成统计城市同簇次数的对称矩阵的实现方案
城市聚类共簇次数统计方案
问题说明
我有一个按年份划分的城市聚类数据库,该数据库是基于**模块度(modularity)对不同年份的城市数据应用社区检测(community detection)**算法得到的。示例数据如下:
v1 city cluster year 0 "city1" 0 2000 1 "city2" 2 2000 2 "city3" 1 2000 3 "city4" 0 2000 4 "city5" 2 2000 0 "city1" 2 2001 1 "city2" 1 2001 2 "city3" 0 2001 3 "city4" 0 2001 4 "city5" 0 2001 0 "city1" 1 2002 1 "city2" 2 2002 2 "city3" 0 2002 3 "city4" 0 2002 4 "city5" 1 2002
需要统计每对城市每年同属一个簇的次数,最终得到一个对称矩阵,矩阵的行和列均为城市,每个元素代表对应城市对在所有年份中同属同一簇的次数(与具体簇编号无关)。示例结果如下:
city1 city2 city3 city4 city5 city1 . 0 0 1 1 city2 0 . 0 0 1 city3 0 0 . 2 1 city4 1 0 2 . 1 city5 1 1 1 1 .
采用Python开发,也接受Matlab或R语言的实现方案。
Python实现方案
代码实现
import pandas as pd import numpy as np # 读取数据(实际场景可替换为pd.read_csv等文件读取逻辑) data = pd.DataFrame({ 'city': ['city1', 'city2', 'city3', 'city4', 'city5']*3, 'cluster': [0, 2, 1, 0, 2, 2, 1, 0, 0, 0, 1, 2, 0, 0, 1], 'year': [2000]*5 + [2001]*5 + [2002]*5 }) # 获取排序后的唯一城市列表,保证矩阵顺序一致 cities = sorted(data['city'].unique()) n_cities = len(cities) # 初始化计数矩阵 count_matrix = np.zeros((n_cities, n_cities)) # 按年份分组处理每一年的聚类结果 for year, group in data.groupby('year'): city_to_cluster = dict(zip(group['city'], group['cluster'])) # 遍历所有不重复的城市对 for i in range(n_cities): for j in range(i+1, n_cities): city_a, city_b = cities[i], cities[j] if city_to_cluster[city_a] == city_to_cluster[city_b]: count_matrix[i][j] += 1 count_matrix[j][i] += 1 # 转为DataFrame并设置行列索引,对角线标记为无意义的'.' result_df = pd.DataFrame(count_matrix, index=cities, columns=cities) np.fill_diagonal(result_df.values, '.') print(result_df)
R语言实现方案
代码实现
# 构造示例数据(实际场景可替换为read.csv等读取逻辑) data <- data.frame( city = rep(c("city1", "city2", "city3", "city4", "city5"), 3), cluster = c(0,2,1,0,2, 2,1,0,0,0, 1,2,0,0,1), year = rep(c(2000,2001,2002), each=5) ) # 获取唯一城市列表 cities <- unique(data$city) n_cities <- length(cities) # 初始化计数矩阵并设置行列名称 count_matrix <- matrix(0, nrow=n_cities, ncol=n_cities, dimnames=list(cities, cities)) # 按年份遍历处理 years <- unique(data$year) for (y in years) { year_data <- subset(data, year == y) city_to_cluster <- setNames(year_data$cluster, year_data$city) # 遍历所有不重复的城市对 for (i in 1:(n_cities-1)) { for (j in (i+1):n_cities) { if (city_to_cluster[cities[i]] == city_to_cluster[cities[j]]) { count_matrix[i,j] <- count_matrix[i,j] + 1 count_matrix[j,i] <- count_matrix[j,i] + 1 } } } } # 对角线替换为'.' diag(count_matrix) <- "." print(count_matrix)
Matlab实现方案
代码实现
% 构造示例数据(实际场景可替换为readtable等读取逻辑) city = repmat({'city1','city2','city3','city4','city5'}, 1, 3); cluster = [0,2,1,0,2, 2,1,0,0,0, 1,2,0,0,1]; year = repmat([2000,2001,2002], 1, 5); data = table(city', cluster', year', 'VariableNames', {'city','cluster','year'}); % 获取唯一城市列表 cities = unique(data.city); n_cities = length(cities); % 初始化计数矩阵 count_matrix = zeros(n_cities); % 按年份遍历处理 years = unique(data.year); for y = years year_data = data(data.year == y, :); % 创建城市到簇的映射 city_to_cluster = containers.Map(year_data.city, year_data.cluster); % 遍历所有不重复的城市对 for i = 1:n_cities-1 for j = i+1:n_cities if city_to_cluster(cities{i}) == city_to_cluster(cities{j}) count_matrix(i,j) = count_matrix(i,j) + 1; count_matrix(j,i) = count_matrix(j,i) + 1; end end end end % 转换为表格并设置行列名称,对角线替换为'.' count_matrix = array2table(count_matrix, 'RowNames', cities, 'VariableNames', cities); count_matrix(1:n_cities+1:end) = {'.'}; disp(count_matrix);
内容的提问来源于stack exchange,提问作者Lusian
相关产品推荐
相关产品推荐

