You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Polars API高效生成带负采样的推荐系统用户历史

用Polars高效生成带负采样的用户历史推荐数据

问题背景

我正在开发一款推荐系统,需要借助Polars API高效生成带负采样的用户历史数据。现有两个数据集:

1. 用户-文章交互数据集

包含用户已读文章的user_id和article_id,示例代码:

import polars as pl
    
df_user_articles = pl.DataFrame({
    'user_id': [1, 1, 2, 2, 2, 3, 4, 4, 4, 4],
    'article_id': [101, 102, 103, 104, 105, 106, 107, 108, 109, 110]
})

2. 文章数据集

包含所有article_id及对应文本数据,示例代码:

df_articles = pl.DataFrame({
    'article_id': list(range(100, 200)),  # 示例文章ID
    'text': ['text'] * 100
})

需求

  • 为每个user_id生成长度在5-20之间的已读文章历史列表;
  • 为每个用户生成候选样本:已读文章标记为正样本(后缀-1),未读文章标记为负样本(后缀-0);
  • 最终输出包含user_id、user_history、candidates列的DataFrame,格式示例:
user_iduser_historycandidates
1['101', '102']['101-1', '102-1', '103-0']
2['103', '104', '105']['103-1', '104-1', '105-1', '106-0']

约束条件

  • 避免使用apply方法,保证大规模数据集处理效率;
  • 候选生成时无需校验文章是否被用户阅读。

现有问题

当前方案无法保留article_id的字符串类型,且未满足全部需求,代码如下:

matrix_size = users_df.shape[0]

num_candidates = 5

index_matrix = np.random.randint(0, articles_df.shape[0], size=(matrix_size, num_candidates))

users_df.with_columns(
    candidates=articles_df['article_id'].to_numpy()[:,np.newaxis][index_matrix].reshape(matrix_size, num_candidates)
)

解决方案

步骤1:生成用户历史列表

按用户分组聚合已读文章,统一调整列表长度到5-20范围内:

# 生成用户历史,处理长度在5-20之间
df_user_history = df_user_articles.group_by('user_id').agg(
    pl.col('article_id').cast(str).alias('raw_history')
).with_columns(
    # 长度不足5则重复填充,超过20则截断
    user_history=pl.when(pl.col('raw_history').list.len() < 5)
        .then(pl.col('raw_history').list.repeat(5).list.slice(0, 5))
        .when(pl.col('raw_history').list.len() > 20)
        .then(pl.col('raw_history').list.slice(0, 20))
        .otherwise(pl.col('raw_history'))
).drop('raw_history')

步骤2:生成正负候选样本

用Polars矢量化操作生成正、负样本,合并后得到最终候选列表:

# 全局文章ID列表(转字符串)
all_articles = df_articles.select(pl.col('article_id').cast(str)).to_series()

# 自定义负样本数量
num_neg_samples = 5

# 生成候选样本
df_result = df_user_history.with_columns(
    # 正样本:已读文章拼接"-1"
    pos_candidates=pl.col('user_history').list.eval(pl.element() + "-1"),
    # 负样本:随机采样全局文章拼接"-0",矢量化操作避免apply
    neg_candidates=pl.int_rand(0, len(all_articles), num_neg_samples)
        .list.map(lambda idx: all_articles[idx] + "-0")
).with_columns(
    # 合并正负候选为一个列表
    candidates=pl.col('pos_candidates') + pl.col('neg_candidates')
).drop('pos_candidates', 'neg_candidates')

完整代码

import polars as pl

# 初始化数据集
df_user_articles = pl.DataFrame({
    'user_id': [1, 1, 2, 2, 2, 3, 4, 4, 4, 4],
    'article_id': [101, 102, 103, 104, 105, 106, 107, 108, 109, 110]
})

df_articles = pl.DataFrame({
    'article_id': list(range(100, 200)),
    'text': ['text'] * 100
})

# 生成用户历史
df_user_history = df_user_articles.group_by('user_id').agg(
    pl.col('article_id').cast(str).alias('raw_history')
).with_columns(
    user_history=pl.when(pl.col('raw_history').list.len() < 5)
        .then(pl.col('raw_history').list.repeat(5).list.slice(0, 5))
        .when(pl.col('raw_history').list.len() > 20)
        .then(pl.col('raw_history').list.slice(0, 20))
        .otherwise(pl.col('raw_history'))
).drop('raw_history')

# 生成候选样本
all_articles = df_articles.select(pl.col('article_id').cast(str)).to_series()
num_neg_samples = 5

df_result = df_user_history.with_columns(
    pos_candidates=pl.col('user_history').list.eval(pl.element() + "-1"),
    neg_candidates=pl.int_rand(0, len(all_articles), num_neg_samples)
        .list.map(lambda idx: all_articles[idx] + "-0")
).with_columns(
    candidates=pl.col('pos_candidates') + pl.col('neg_candidates')
).drop('pos_candidates', 'neg_candidates')

print(df_result)

关键说明

  • 全程使用Polars矢量化操作,无apply调用,适配大规模数据处理;
  • article_id通过cast(str)转为字符串类型,完全匹配示例格式;
  • 负采样用pl.int_rand生成随机索引,直接映射全局文章列表,无需校验用户已读状态;
  • 用户历史长度处理逻辑可按需调整(比如不足5时填充默认文章而非重复)。

内容的提问来源于stack exchange,提问作者MPA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 11:33:22