如何通过Do Loop高效实现跨数据集关键词匹配的商品标签列创建?
解决方案:使用循环自动生成商品标签列
我们可以通过循环遍历商品类别来自动生成每个ID的标签列,避免手动编写大量if-then语句。以下是几种主流编程语言的实现方式:
1. SAS实现(宏循环+数据步循环)
步骤1:准备数据
/* 创建通用商品列表数据集 */ data goods_list; input Keyword $ Goods $; datalines; Soap A Soap Soap B Soap Shampoo Shampoo ; run; /* 创建交易数据集 */ data transactions; input ID Date $ Txn $; datalines; 1 1/22 "Soap A 100 ml" 1 1/23 "Soap A 50 ml" 2 1/24 "Soap B 100 ml" 2 1/24 "Shampoo 50 ml" 3 1/24 "Juice 100 g" ; run;
步骤2:自动生成标签列
/* 获取唯一商品类别并存储为宏变量 */ proc sql noprint; select distinct Goods into :goods_list separated by ' ' from goods_list; select count(distinct Goods) into :num_goods from goods_list; quit; /* 循环生成标签并匹配交易记录 */ data id_goods_flags; set transactions; by ID; /* 每个ID首次出现时初始化所有标签为0 */ if first.ID then do; %do i=1 %to &num_goods; %let current_good = %scan(&goods_list, &i); ¤t_good = 0; %end; end; /* 哈希表快速映射关键词到商品类别 */ if _N_ = 1 then do; declare hash goods_map(dataset:'goods_list'); goods_map.defineKey('Keyword'); goods_map.defineData('Goods'); goods_map.defineDone(); call missing(Keyword, Goods); end; /* 遍历所有关键词,检查当前交易是否匹配 */ do _iter = 1 to goods_map.num_items(); goods_map.next(); if index(Txn, Keyword) then do; call symputx('matched_good', Goods); &matched_good = 1; end; end; /* 仅保留每个ID的最终标签记录 */ if last.ID then output; keep ID &goods_list; run; /* 查看结果 */ proc print data=id_goods_flags; run;
2. R语言实现(嵌套for循环)
步骤1:准备数据
goods_list <- data.frame( Keyword = c("Soap A", "Soap B", "Shampoo"), Goods = c("Soap", "Soap", "Shampoo"), stringsAsFactors = FALSE ) transactions <- data.frame( ID = c(1, 1, 2, 2, 3), Date = c("1/22", "1/23", "1/24", "1/24", "1/24"), Txn = c("Soap A 100 ml", "Soap A 50 ml", "Soap B 100 ml", "Shampoo 50 ml", "Juice 100 g"), stringsAsFactors = FALSE )
步骤2:循环生成标签列
# 获取唯一商品类别 unique_goods <- unique(goods_list$Goods) # 初始化结果数据框,标签默认值为0 result <- data.frame(ID = unique(transactions$ID)) result[, unique_goods] <- 0 # 遍历每个ID for (id in result$ID) { user_txns <- transactions$Txn[transactions$ID == id] # 遍历每个商品类别 for (good in unique_goods) { target_keywords <- goods_list$Keyword[goods_list$Goods == good] # 检查是否有交易匹配任意关键词 has_match <- any(sapply(target_keywords, function(k) grepl(k, user_txns))) if (has_match) { result[result$ID == id, good] <- 1 } } } # 输出结果 print(result)
3. Python实现(Pandas+循环)
步骤1:准备数据
import pandas as pd goods_list = pd.DataFrame({ 'Keyword': ['Soap A', 'Soap B', 'Shampoo'], 'Goods': ['Soap', 'Soap', 'Shampoo'] }) transactions = pd.DataFrame({ 'ID': [1, 1, 2, 2, 3], 'Date': ['1/22', '1/23', '1/24', '1/24', '1/24'], 'Txn': ['Soap A 100 ml', 'Soap A 50 ml', 'Soap B 100 ml', 'Shampoo 50 ml', 'Juice 100 g'] })
步骤2:循环生成标签列
# 获取唯一商品类别 unique_goods = goods_list['Goods'].unique() # 初始化结果数据框 result = pd.DataFrame({'ID': transactions['ID'].unique()}) for good in unique_goods: result[good] = 0 # 遍历每个ID的交易记录 for idx, row in result.iterrows(): current_id = row['ID'] txns = transactions[transactions['ID'] == current_id]['Txn'] # 检查每个商品类别 for good in unique_goods: keywords = goods_list[goods_list['Goods'] == good]['Keyword'] match_found = any(txn.str.contains(keyword).any() for keyword in keywords) if match_found: result.at[idx, good] = 1 # 打印结果 print(result)
所有方案都会输出你需要的结果:
| ID | Soap | Shampoo |
|---|---|---|
| 1 | 1 | 0 |
| 2 | 1 | 1 |
| 3 | 0 | 0 |
内容的提问来源于stack exchange,提问作者user23560499
相关产品推荐
相关产品推荐

