You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中从商品描述字符串提取特定商品类型文本?

从商品描述中提取商品类型的Python实现方法

根据你的需求,这里提供几种实用的实现方案,适用于不同场景:

一、规则匹配法(适合已知商品类型集合的场景)

如果你的商品类型是固定且可枚举的(比如示例中的saree、lehenga、swim suit),直接用关键词匹配是最简单高效的方式:

  • 先定义好商品类型集合,遍历每条描述时检查关键词是否存在,同时处理大小写不敏感问题
  • 代码示例:
import pandas as pd

# 初始化示例数据
data = {
    'product_description': [
        'kanchivaram saree of red colour',
        'Pink gujrati saree',
        'Lehenga from Surat',
        'Red swim suit'
    ]
}
df = pd.DataFrame(data)

# 已知商品类型集合
product_types = {'saree', 'lehenga', 'swim suit'}

def extract_type(desc):
    desc_lower = desc.lower()
    for p_type in product_types:
        if p_type in desc_lower:
            return p_type
    return None  # 无匹配时返回空值

df['product_type'] = df['product_description'].apply(extract_type)
print(df)

这种方法维护成本低,但需要手动更新类型集合来覆盖新的商品类型。

二、NLP词性标注法(适合灵活提取未知商品类型的场景)

如果需要处理未知的商品类型,可以借助NLP工具提取描述中的核心名词/名词短语:

  • 使用spaCy或nltk这类库做词性标注,筛选出描述中的核心名词短语
  • 代码示例(基于spaCy):
import pandas as pd
import spacy

# 加载英文NLP模型
nlp = spacy.load("en_core_web_sm")

data = {
    'product_description': [
        'kanchivaram saree of red colour',
        'Pink gujrati saree',
        'Lehenga from Surat',
        'Red swim suit'
    ]
}
df = pd.DataFrame(data)

def extract_type(desc):
    doc = nlp(desc)
    # 提取名词短语,过滤掉颜色、产地这类非商品类型的名词
    noun_phrases = [chunk.text.lower() for chunk in doc.noun_chunks 
                   if chunk.root.pos_ in ('NOUN', 'PROPN') 
                   and chunk.text.lower() not in {'red', 'pink', 'surat', 'colour'}]
    # 优先匹配已知类型,无匹配时返回最长的核心名词短语
    for phrase in noun_phrases:
        if phrase in {'saree', 'lehenga', 'swim suit'}:
            return phrase
    return max(noun_phrases, key=len) if noun_phrases else None

df['product_type'] = df['product_description'].apply(extract_type)
print(df)

这种方法灵活性更强,但需要根据实际场景调整过滤规则,避免提取无关内容。

三、监督学习法(适合有大量标注数据的场景)

如果有足够多的标注数据(如示例中的对应关系),可以训练分类模型来自动识别商品类型:

  • 用scikit-learn构建文本分类管道,或者用BERT等深度学习模型做序列标注
  • 代码示例(基于scikit-learn):
import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.svm import SVC
from sklearn.pipeline import Pipeline

# 标注好的训练数据
train_data = {
    'product_description': [
        'kanchivaram saree of red colour',
        'Pink gujrati saree',
        'Lehenga from Surat',
        'Red swim suit',
        'Blue cotton saree',
        'Designer lehenga'
    ],
    'product_type': ['saree', 'saree', 'lehenga', 'swim suit', 'saree', 'lehenga']
}
train_df = pd.DataFrame(train_data)

# 构建TF-IDF+SVM分类管道
pipeline = Pipeline([
    ('tfidf', TfidfVectorizer(stop_words='english')),
    ('classifier', SVC(kernel='linear'))
])

# 训练模型
pipeline.fit(train_df['product_description'], train_df['product_type'])

# 预测新数据
test_data = {'product_description': ['Green silk saree', 'Black swim suit']}
test_df = pd.DataFrame(test_data)
test_df['product_type'] = pipeline.predict(test_df['product_description'])
print(test_df)

这种方法准确率更高,能处理复杂的描述文本,但需要足够的标注数据支撑模型训练。

内容的提问来源于stack exchange,提问作者Aditya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 16:24:30