基于Polars API高效生成带负采样的推荐系统用户历史
用Polars高效生成带负采样的用户历史推荐数据
问题背景
我正在开发一款推荐系统,需要借助Polars API高效生成带负采样的用户历史数据。现有两个数据集:
1. 用户-文章交互数据集
包含用户已读文章的user_id和article_id,示例代码:
import polars as pl df_user_articles = pl.DataFrame({ 'user_id': [1, 1, 2, 2, 2, 3, 4, 4, 4, 4], 'article_id': [101, 102, 103, 104, 105, 106, 107, 108, 109, 110] })
2. 文章数据集
包含所有article_id及对应文本数据,示例代码:
df_articles = pl.DataFrame({ 'article_id': list(range(100, 200)), # 示例文章ID 'text': ['text'] * 100 })
需求
- 为每个
user_id生成长度在5-20之间的已读文章历史列表; - 为每个用户生成候选样本:已读文章标记为正样本(后缀
-1),未读文章标记为负样本(后缀-0); - 最终输出包含
user_id、user_history、candidates列的DataFrame,格式示例:
| user_id | user_history | candidates |
|---|---|---|
| 1 | ['101', '102'] | ['101-1', '102-1', '103-0'] |
| 2 | ['103', '104', '105'] | ['103-1', '104-1', '105-1', '106-0'] |
约束条件
- 避免使用
apply方法,保证大规模数据集处理效率; - 候选生成时无需校验文章是否被用户阅读。
现有问题
当前方案无法保留article_id的字符串类型,且未满足全部需求,代码如下:
matrix_size = users_df.shape[0] num_candidates = 5 index_matrix = np.random.randint(0, articles_df.shape[0], size=(matrix_size, num_candidates)) users_df.with_columns( candidates=articles_df['article_id'].to_numpy()[:,np.newaxis][index_matrix].reshape(matrix_size, num_candidates) )
解决方案
步骤1:生成用户历史列表
按用户分组聚合已读文章,统一调整列表长度到5-20范围内:
# 生成用户历史,处理长度在5-20之间 df_user_history = df_user_articles.group_by('user_id').agg( pl.col('article_id').cast(str).alias('raw_history') ).with_columns( # 长度不足5则重复填充,超过20则截断 user_history=pl.when(pl.col('raw_history').list.len() < 5) .then(pl.col('raw_history').list.repeat(5).list.slice(0, 5)) .when(pl.col('raw_history').list.len() > 20) .then(pl.col('raw_history').list.slice(0, 20)) .otherwise(pl.col('raw_history')) ).drop('raw_history')
步骤2:生成正负候选样本
用Polars矢量化操作生成正、负样本,合并后得到最终候选列表:
# 全局文章ID列表(转字符串) all_articles = df_articles.select(pl.col('article_id').cast(str)).to_series() # 自定义负样本数量 num_neg_samples = 5 # 生成候选样本 df_result = df_user_history.with_columns( # 正样本:已读文章拼接"-1" pos_candidates=pl.col('user_history').list.eval(pl.element() + "-1"), # 负样本:随机采样全局文章拼接"-0",矢量化操作避免apply neg_candidates=pl.int_rand(0, len(all_articles), num_neg_samples) .list.map(lambda idx: all_articles[idx] + "-0") ).with_columns( # 合并正负候选为一个列表 candidates=pl.col('pos_candidates') + pl.col('neg_candidates') ).drop('pos_candidates', 'neg_candidates')
完整代码
import polars as pl # 初始化数据集 df_user_articles = pl.DataFrame({ 'user_id': [1, 1, 2, 2, 2, 3, 4, 4, 4, 4], 'article_id': [101, 102, 103, 104, 105, 106, 107, 108, 109, 110] }) df_articles = pl.DataFrame({ 'article_id': list(range(100, 200)), 'text': ['text'] * 100 }) # 生成用户历史 df_user_history = df_user_articles.group_by('user_id').agg( pl.col('article_id').cast(str).alias('raw_history') ).with_columns( user_history=pl.when(pl.col('raw_history').list.len() < 5) .then(pl.col('raw_history').list.repeat(5).list.slice(0, 5)) .when(pl.col('raw_history').list.len() > 20) .then(pl.col('raw_history').list.slice(0, 20)) .otherwise(pl.col('raw_history')) ).drop('raw_history') # 生成候选样本 all_articles = df_articles.select(pl.col('article_id').cast(str)).to_series() num_neg_samples = 5 df_result = df_user_history.with_columns( pos_candidates=pl.col('user_history').list.eval(pl.element() + "-1"), neg_candidates=pl.int_rand(0, len(all_articles), num_neg_samples) .list.map(lambda idx: all_articles[idx] + "-0") ).with_columns( candidates=pl.col('pos_candidates') + pl.col('neg_candidates') ).drop('pos_candidates', 'neg_candidates') print(df_result)
关键说明
- 全程使用Polars矢量化操作,无
apply调用,适配大规模数据处理; article_id通过cast(str)转为字符串类型,完全匹配示例格式;- 负采样用
pl.int_rand生成随机索引,直接映射全局文章列表,无需校验用户已读状态; - 用户历史长度处理逻辑可按需调整(比如不足5时填充默认文章而非重复)。
内容的提问来源于stack exchange,提问作者MPA
相关产品推荐
相关产品推荐

