在Memgraph中创建查询模块:统计音乐流派Top N位置出现占比
解决Memgraph中音乐流派Top N偏好百分比统计问题
一、纯Cypher查询方案
如果只是临时统计,直接用Cypher就能实现,支持自定义Top N值,同时输出两种常用的百分比维度:
查询代码
// 自定义Top N数值,这里设为3,可按需修改 WITH 3 AS top_n MATCH (u) WHERE size(u.genres) > 0 // 过滤无流派偏好的用户 UNWIND range(0, top_n - 1) AS idx WHERE idx < size(u.genres) // 避免用户流派数不足Top N时数组越界 WITH top_n, u.genres[idx] AS genre, count(*) AS total_occurrences, count(DISTINCT u) AS unique_users // 计算所有用户的Top N位置总数量 WITH top_n, genre, total_occurrences, unique_users, sum(least(size(u.genres), top_n)) OVER() AS total_top_positions, count(DISTINCT u) OVER() AS total_valid_users // 计算两种百分比并排序 RETURN genre, round((total_occurrences / total_top_positions) * 100, 2) AS 占TopN位置百分比, round((unique_users / total_valid_users) * 100, 2) AS 含该流派TopN的用户占比 ORDER BY 占TopN位置百分比 DESC;
关键说明
- 占TopN位置百分比:该流派在所有用户的Top N偏好位置中,出现次数占总Top N位置数的比例
- 含该流派TopN的用户占比:有多少比例的用户把该流派放在了自己的Top N偏好里
- 可以通过修改
WITH 3 AS top_n中的数值来切换不同的Top N统计,比如改成5就是统计Top5偏好
二、Query Module方案(Python)
如果需要重复调用统计逻辑,或者后续要扩展更复杂的分析(比如结合好友关系联动统计),可以写一个Python Query Module:
模块代码
创建genre_stats.py文件,放到Memgraph的query_modules目录(默认路径:/var/lib/memgraph/query_modules/):
import mgp @mgp.read_proc def top_n_genre_percentages(ctx: mgp.ProcCtx, top_n: int) -> mgp.Record(genre=str, position_percentage=float, user_percentage=float): total_users = 0 genre_occurrences = {} genre_user_set = {} total_top_positions = 0 # 遍历所有用户节点 for user in ctx.graph.nodes: genres = user.properties.get("genres", []) if not genres: continue total_users += 1 # 提取用户的Top N流派 top_genres = genres[:top_n] total_top_positions += len(top_genres) for genre in top_genres: # 统计流派出现次数 genre_occurrences[genre] = genre_occurrences.get(genre, 0) + 1 # 统计有多少不同用户包含该流派在Top N中 if genre not in genre_user_set: genre_user_set[genre] = set() genre_user_set[genre].add(user.id) # 生成结果并排序 results = [] for genre in genre_occurrences: pos_percent = (genre_occurrences[genre] / total_top_positions) * 100 user_percent = (len(genre_user_set[genre]) / total_users) * 100 results.append(mgp.Record( genre=genre, position_percentage=round(pos_percent, 2), user_percentage=round(user_percent, 2) )) results.sort(key=lambda x: x.position_percentage, reverse=True) return results
使用方法
- 加载模块:在Memgraph Lab中执行
CALL mg.load_all(); - 调用统计:比如统计Top3偏好,执行
CALL genre_stats.top_n_genre_percentages(3) YIELD *;
注意事项
- 两种方案都会自动过滤没有流派偏好的用户,避免影响统计准确性
- 百分比保留两位小数,便于阅读和对比
- 如果你的数据集里用户流派列表是字符串格式(不是数组),需要先通过
split()函数转成数组,比如split(u.genres, ",")(假设用逗号分隔)
内容的提问来源于stack exchange,提问作者mattrixxxx
相关产品推荐
相关产品推荐

