如何为商品描述变体拆分添加匹配词数量过滤规则?
问题描述
我正在做商品描述变体拆分的分组工作,目前已经实现了将商品描述排序后对比最相似相邻项并移除差异的功能。现在需要添加过滤规则:仅当两个描述的匹配词数量至少为n(示例中n取3)时,才移除差异。
输入示例
| Product Description |
|---|
| Petzl Red HMS Carabiner |
| Petzl Blue HMS Carabiner |
| Petzl HMS Carabiner Orange |
| Petzl Green Carabiner |
| Petzl Purple Carabiner |
| Liquid Chalk - 100ml |
| Liquid Chalk - 100ml (Case of 10) |
期望输出
| Product Description | Variance |
|---|---|
| Petzl HMS Carabiner | Red |
| Petzl HMS Carabiner | Blue |
| Petzl HMS Carabiner | Orange |
| Petzl Green Carabiner | NaN |
| Petzl Purple Carabiner | NaN |
| Liquid Chalk - 100ml | NaN |
| Liquid Chalk - 100ml | (Case of 10) |
当前未添加过滤规则的代码
import pandas as pd def get_intersection(descr1, descr2): if pd.isna(descr1) or pd.isna(descr2): return set() return set(descr1.split()).intersection(set(descr2.split())) def get_unique_words(descr, intersection): unique_words = " ".join( word for word in descr.split() if word not in intersection ) if len(unique_words) > 0: return unique_words def get_unique_description(row): if len(row["next_product_intersection"]) == 0 and len(row["prev_product_intersection"]) == 0: return row["Product Description"] if len(row["next_product_intersection"]) >= len(row["prev_product_intersection"]): return row["next_product_unique_words"] return row["prev_product_unique_words"] df = pd.DataFrame([ "Petzl Red HMS Carabiner", "Petzl Blue HMS Carabiner", "Petzl HMS Carabiner Orange", "Petzl Green Carabiner", "Petzl Purple Carabiner", "Liquid Chalk - 100ml", "Liquid Chalk - 100ml (Case of 10)" ], columns=["Product Description"]) df["next_product"] = df["Product Description"].shift(-1) df["prev_product"] = df["Product Description"].shift(1) df["next_product_intersection"] = df.apply( lambda row: get_intersection(row["Product Description"], row["next_product"]), axis=1 ) df["prev_product_intersection"] = df.apply( lambda row: get_intersection(row["Product Description"], row["prev_product"]), axis=1 ) df["next_product_unique_words"] = df.apply( lambda row: get_unique_words(row["Product Description"], row["next_product_intersection"]), axis=1 ) df["prev_product_unique_words"] = df.apply( lambda row: get_unique_words(row["Product Description"], row["prev_product_intersection"]), axis=1 ) df["Variance"] = df.apply(get_unique_description, axis=1) df = df[["Product Description", "Variance"]] print(df)
解决方案
要整合匹配词数量过滤规则,核心是在判断是否移除差异前,先检查相邻描述的匹配词数量是否达到阈值n(这里n=3)。具体修改步骤如下:
- 定义匹配词阈值:设定全局变量统一控制阈值,方便后续调整。
- 修改差异提取逻辑:在选择使用相邻项的差异前,先判断对应交集的大小是否达标;未达标则不拆分差异。
- 生成统一基准描述:当差异有效时,通过原描述减去差异词得到标准化的基准描述,匹配期望输出格式。
修改后的完整代码
import pandas as pd # 定义匹配词数量阈值 MATCH_THRESHOLD = 3 def get_intersection(descr1, descr2): if pd.isna(descr1) or pd.isna(descr2): return set() return set(descr1.split()).intersection(set(descr2.split())) def get_unique_words(descr, intersection): unique_words = " ".join( word for word in descr.split() if word not in intersection ) return unique_words if len(unique_words) > 0 else None def get_unique_description(row): next_inter_len = len(row["next_product_intersection"]) prev_inter_len = len(row["prev_product_intersection"]) # 前后交集都未达标,不拆分差异 if next_inter_len < MATCH_THRESHOLD and prev_inter_len < MATCH_THRESHOLD: return None # 选择达标且交集更大的一侧提取差异 use_next = next_inter_len >= prev_inter_len and next_inter_len >= MATCH_THRESHOLD use_prev = prev_inter_len > next_inter_len and prev_inter_len >= MATCH_THRESHOLD if use_next: return row["next_product_unique_words"] elif use_prev: return row["prev_product_unique_words"] else: return None def get_base_description(row): if pd.notna(row["Variance"]): # 原描述去掉差异词,生成统一基准描述 base_words = [word for word in row["Product Description"].split() if word not in row["Variance"].split()] return " ".join(base_words) else: return row["Product Description"] # 构造示例数据 df = pd.DataFrame([ "Petzl Red HMS Carabiner", "Petzl Blue HMS Carabiner", "Petzl HMS Carabiner Orange", "Petzl Green Carabiner", "Petzl Purple Carabiner", "Liquid Chalk - 100ml", "Liquid Chalk - 100ml (Case of 10)" ], columns=["Product Description"]) df["next_product"] = df["Product Description"].shift(-1) df["prev_product"] = df["Product Description"].shift(1) df["next_product_intersection"] = df.apply( lambda row: get_intersection(row["Product Description"], row["next_product"]), axis=1 ) df["prev_product_intersection"] = df.apply( lambda row: get_intersection(row["Product Description"], row["prev_product"]), axis=1 ) df["next_product_unique_words"] = df.apply( lambda row: get_unique_words(row["Product Description"], row["next_product_intersection"]), axis=1 ) df["prev_product_unique_words"] = df.apply( lambda row: get_unique_words(row["Product Description"], row["prev_product_intersection"]), axis=1 ) # 计算差异列 df["Variance"] = df.apply(get_unique_description, axis=1) # 生成基准描述列 df["Product Description"] = df.apply(get_base_description, axis=1) # 保留目标列并输出 df = df[["Product Description", "Variance"]] print(df)
关键改动说明
- 阈值控制:通过
MATCH_THRESHOLD统一设定匹配词数量要求,后续可根据业务需求灵活调整数值。 - 逻辑判断优化:在
get_unique_description中,只有相邻项的交集长度达标时才提取差异,否则返回None对应最终的NaN。 - 基准描述生成:新增
get_base_description函数,将有效差异对应的原描述转换为统一的基准文本,完全匹配期望输出格式。
内容的提问来源于stack exchange,提问作者hydroproxy
相关产品推荐
相关产品推荐

