TraMineR序列分析:批量计算距离是否可行?资源优化咨询
APP使用序列分析的资源优化与方法疑问
数据集与分析目标
- 身份:传播学者,TraMineR与序列分析新手
- 数据集结构:
| 参与者ID | 会话ID | 使用的APP类别 | 开始时间(Unix时间) | 结束时间(Unix时间) |
|---|---|---|---|---|
| 0001 | 0001_1 | Communication | 1614868224 | 1614868236 |
| 0001 | 0001_1 | Social Media | 1614868236 | 1614868265 |
| 0002 | 0002_1 | Games | 1614868265 | 1614868320 |
| ... | ... | ... | ... | ... |
- 分析层面:
- 参与者层面
- 会话层面(智能手机屏幕点亮至熄灭的连贯使用序列)
- 数据规模:近400名参与者,每人2000-5000个会话,总数据约140万个会话
- 核心目标:识别连续使用的APP类别序列
当前代码与资源瓶颈
以下是测试代码:
labels = seqstatl(sample$app_category) states = 1:length(labels) session_seq = seqdef(data = sample, var = c("session", "begin", "end", "app_category"), informat = "SPELL", states = states, labels = labels, process = FALSE) print(session_seq[1:15, ], format = "SPS") # Using the transition rates between states observed in the sequence data cost = seqsubm(session_seq, method = "TRATE", with.missing = TRUE) # compute the distances using the matrix and the default indel cost of 1 session_seq_OM = seqdist(session_seq, method = "OM", sm = cost, with.missing = TRUE) # --> Function crashed due to lack of RAM
- 遇到的问题:运行
seqdist计算OM距离时因内存不足崩溃;即使使用单参与者约4000个会话生成的STS序列(含530个对象、1222844个变量),在1TB内存机器上仍因内存不足终止计算
核心疑问
- 有没有方法提升计算资源使用效率,避免内存溢出?
- 能否拆分数据集批量计算序列距离后再合并?这种操作会导致分析结果失真吗?
- 除TraMineR官方用户指南外,还有哪些实用的参考资料可以学习?
内容的提问来源于stack exchange,提问作者confused_person
相关产品推荐
相关产品推荐

