You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python NLP库缩短DataFrame中产品标题至40字符以内

处理DataFrame中过长产品名称的解决方案

针对产品名称过长(超过40字符)需要生成有意义短名称的需求,不能直接截断(会导致语义缺失),我们可以通过提取关键信息并重组的方式实现。核心思路是保留产品核心属性、尺寸、颜色等关键信息,剔除无关的包装/数量信息,再重组为符合长度要求的有效名称。

示例场景

输入:Cut Resistant Gloves, Size 8, Grey/Black - 12 per DZ(52字符)
期望输出:Resistant Size 8 Grey/Black Gloves(34字符)

代码实现

import pandas as pd

def shorten_product_name(name, max_len=40):
    # 移除末尾的数量/包装信息
    name = name.split(' - ')[0]
    # 拆分名称为属性片段
    segments = [seg.strip() for seg in name.split(',')]
    
    key_elements = []
    product_core = ''
    # 定义常见产品核心名词(可根据业务扩展)
    core_products = {'gloves', 'boots', 't-shirt', 'hat', 'pants'}
    
    for seg in segments:
        words = seg.lower().split()
        # 提取尺寸信息
        if 'size' in words:
            key_elements.append(seg)
        # 提取颜色信息(通常含/或颜色词)
        elif any(char in seg for char in ['/', 'white', 'black', 'grey', 'brown']):
            key_elements.append(seg)
        else:
            # 提取属性并识别核心产品
            for word in seg.split():
                if word.lower() in core_products:
                    product_core = word
                else:
                    key_elements.append(word)
    
    # 把核心产品名放在最后,保证语义完整
    if product_core:
        key_elements.append(product_core)
    
    # 生成初始短名称,若超长则逐步移除非关键属性
    short_name = ' '.join(key_elements)
    while len(short_name) > max_len and key_elements:
        # 优先移除非尺寸、非颜色的属性
        for idx, elem in enumerate(key_elements):
            if not elem.startswith('Size') and '/' not in elem:
                del key_elements[idx]
                short_name = ' '.join(key_elements)
                break
    
    return short_name

# 测试用DataFrame
df = pd.DataFrame({
    'product_name': [
        'Cut Resistant Gloves, Size 8, Grey/Black - 12 per DZ',
        'Heavy Duty Work Boots, Size 10, Brown Leather - 6 per Case',
        'Premium Cotton T-Shirt, Size M, White - 24 per Box'
    ]
})

# 新增短名称列
df['short_product_name'] = df['product_name'].apply(shorten_product_name)

print(df)

代码说明

  1. 剔除冗余信息:先拆分并移除末尾的数量/包装描述(如- 12 per DZ),减少无效字符占用。
  2. 提取关键元素:优先识别尺寸(含Size)、颜色(含/或常见颜色词),再提取产品属性和核心名称。
  3. 重组与长度控制:将核心产品名放在末尾保证语义,若仍超长则逐步移除非关键属性,直到符合长度要求。
  4. 扩展性:可以根据业务需求扩展core_products集合,或调整关键元素的优先级(比如加入品牌识别)。

内容的提问来源于stack exchange,提问作者jmed1987

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 19:32:37