You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Memgraph中创建查询模块:统计音乐流派Top N位置出现占比

解决Memgraph中音乐流派Top N偏好百分比统计问题

一、纯Cypher查询方案

如果只是临时统计,直接用Cypher就能实现,支持自定义Top N值,同时输出两种常用的百分比维度:

查询代码

// 自定义Top N数值,这里设为3,可按需修改
WITH 3 AS top_n
MATCH (u)
WHERE size(u.genres) > 0  // 过滤无流派偏好的用户
UNWIND range(0, top_n - 1) AS idx
WHERE idx < size(u.genres)  // 避免用户流派数不足Top N时数组越界
WITH top_n, u.genres[idx] AS genre, 
     count(*) AS total_occurrences, 
     count(DISTINCT u) AS unique_users
// 计算所有用户的Top N位置总数量
WITH top_n, genre, total_occurrences, unique_users,
     sum(least(size(u.genres), top_n)) OVER() AS total_top_positions,
     count(DISTINCT u) OVER() AS total_valid_users
// 计算两种百分比并排序
RETURN genre,
       round((total_occurrences / total_top_positions) * 100, 2) AS 占TopN位置百分比,
       round((unique_users / total_valid_users) * 100, 2) AS 含该流派TopN的用户占比
ORDER BY 占TopN位置百分比 DESC;

关键说明

  • 占TopN位置百分比:该流派在所有用户的Top N偏好位置中,出现次数占总Top N位置数的比例
  • 含该流派TopN的用户占比:有多少比例的用户把该流派放在了自己的Top N偏好里
  • 可以通过修改WITH 3 AS top_n中的数值来切换不同的Top N统计,比如改成5就是统计Top5偏好

二、Query Module方案(Python)

如果需要重复调用统计逻辑,或者后续要扩展更复杂的分析(比如结合好友关系联动统计),可以写一个Python Query Module:

模块代码

创建genre_stats.py文件,放到Memgraph的query_modules目录(默认路径:/var/lib/memgraph/query_modules/):

import mgp

@mgp.read_proc
def top_n_genre_percentages(ctx: mgp.ProcCtx, top_n: int) -> mgp.Record(genre=str, position_percentage=float, user_percentage=float):
    total_users = 0
    genre_occurrences = {}
    genre_user_set = {}
    total_top_positions = 0

    # 遍历所有用户节点
    for user in ctx.graph.nodes:
        genres = user.properties.get("genres", [])
        if not genres:
            continue
        total_users += 1
        # 提取用户的Top N流派
        top_genres = genres[:top_n]
        total_top_positions += len(top_genres)
        
        for genre in top_genres:
            # 统计流派出现次数
            genre_occurrences[genre] = genre_occurrences.get(genre, 0) + 1
            # 统计有多少不同用户包含该流派在Top N中
            if genre not in genre_user_set:
                genre_user_set[genre] = set()
            genre_user_set[genre].add(user.id)
    
    # 生成结果并排序
    results = []
    for genre in genre_occurrences:
        pos_percent = (genre_occurrences[genre] / total_top_positions) * 100
        user_percent = (len(genre_user_set[genre]) / total_users) * 100
        results.append(mgp.Record(
            genre=genre,
            position_percentage=round(pos_percent, 2),
            user_percentage=round(user_percent, 2)
        ))
    
    results.sort(key=lambda x: x.position_percentage, reverse=True)
    return results

使用方法

  1. 加载模块:在Memgraph Lab中执行CALL mg.load_all();
  2. 调用统计:比如统计Top3偏好,执行CALL genre_stats.top_n_genre_percentages(3) YIELD *;

注意事项

  • 两种方案都会自动过滤没有流派偏好的用户,避免影响统计准确性
  • 百分比保留两位小数,便于阅读和对比
  • 如果你的数据集里用户流派列表是字符串格式(不是数组),需要先通过split()函数转成数组,比如split(u.genres, ",")(假设用逗号分隔)

内容的提问来源于stack exchange,提问作者mattrixxxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 19:05:36