You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

协作型数据库整理:生成国家合作关联矩阵及汇总变量

Stata处理协作关系表解决方案

数据准备

先将原始数据导入Stata:

clear
input str2 ID str2 country
"A1" "AT"
"A1" "BE"
"A2" "CZ"
"A3" "US"
"A3" "UK"
"A4" "NZ"
end

1. 生成双向计数的协作关系矩阵

该矩阵保留国家间的双向协作记录(如AT-BE和BE-AT各计1次):

preserve
* 标记有协作的ID(包含多个国家)
bysort ID: gen n = _N
keep if n > 1

* 生成同ID下的所有国家配对
bysort ID: expand n
bysort ID: gen match_country = country[_n + _N/_n] if mod(_n, _N/_n) != 0
drop if match_country == "" | match_country == country

* 构建并显示协作矩阵
tab country match_country, matcell(collab_matrix) rowname(colname)
matrix list collab_matrix
restore

运行后得到的矩阵与你期望的第一个结果一致。

2. 生成去重后的单向协作矩阵

仅保留字典序靠前国家的协作记录(如仅保留AT-BE,不重复记录BE-AT):

preserve
bysort ID: gen n = _N
keep if n > 1

bysort ID: expand n
bysort ID: gen match_country = country[_n + _N/_n] if mod(_n, _N/_n) != 0
drop if match_country == "" | match_country == country

* 筛选字典序靠前的配对,实现去重
drop if country > match_country

* 构建并显示去重后的矩阵
tab country match_country, matcell(collab_matrix_unique) rowname(colname)
matrix list collab_matrix_unique
restore

该结果与你期望的去重版矩阵一致。

3. 生成可用于summarize的统计变量

生成每个国家的协作次数变量,以及针对每个国家的二元协作标记变量:

* 标记有协作的ID
bysort ID: gen n = _N

* 生成总协作次数变量
bysort country: gen collab_total = _N - 1 if n > 1
replace collab_total = 0 if collab_total == .

* 生成单个国家的协作标记变量
levelsof country, local(countries)
foreach c of local countries {
    gen collab_with_`c' = (match_country == "`c'") if n > 1
    replace collab_with_`c' = 0 if collab_with_`c' == .
}

* 使用summarize查看统计结果
summarize collab_total collab_with_*

关于你尝试的序号变量优化

你生成的var2是同一ID下的国家序号,可简化为:

bysort ID: gen seq = _n

该变量可用于替代expand方法生成配对,但expand方法更适配多国家的ID场景。

内容的提问来源于stack exchange,提问作者bcast

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 10:12:39