如何在Python中从商品描述字符串提取特定商品类型文本?
从商品描述中提取商品类型的Python实现方法
根据你的需求,这里提供几种实用的实现方案,适用于不同场景:
一、规则匹配法(适合已知商品类型集合的场景)
如果你的商品类型是固定且可枚举的(比如示例中的saree、lehenga、swim suit),直接用关键词匹配是最简单高效的方式:
- 先定义好商品类型集合,遍历每条描述时检查关键词是否存在,同时处理大小写不敏感问题
- 代码示例:
import pandas as pd # 初始化示例数据 data = { 'product_description': [ 'kanchivaram saree of red colour', 'Pink gujrati saree', 'Lehenga from Surat', 'Red swim suit' ] } df = pd.DataFrame(data) # 已知商品类型集合 product_types = {'saree', 'lehenga', 'swim suit'} def extract_type(desc): desc_lower = desc.lower() for p_type in product_types: if p_type in desc_lower: return p_type return None # 无匹配时返回空值 df['product_type'] = df['product_description'].apply(extract_type) print(df)
这种方法维护成本低,但需要手动更新类型集合来覆盖新的商品类型。
二、NLP词性标注法(适合灵活提取未知商品类型的场景)
如果需要处理未知的商品类型,可以借助NLP工具提取描述中的核心名词/名词短语:
- 使用
spaCy或nltk这类库做词性标注,筛选出描述中的核心名词短语 - 代码示例(基于spaCy):
import pandas as pd import spacy # 加载英文NLP模型 nlp = spacy.load("en_core_web_sm") data = { 'product_description': [ 'kanchivaram saree of red colour', 'Pink gujrati saree', 'Lehenga from Surat', 'Red swim suit' ] } df = pd.DataFrame(data) def extract_type(desc): doc = nlp(desc) # 提取名词短语,过滤掉颜色、产地这类非商品类型的名词 noun_phrases = [chunk.text.lower() for chunk in doc.noun_chunks if chunk.root.pos_ in ('NOUN', 'PROPN') and chunk.text.lower() not in {'red', 'pink', 'surat', 'colour'}] # 优先匹配已知类型,无匹配时返回最长的核心名词短语 for phrase in noun_phrases: if phrase in {'saree', 'lehenga', 'swim suit'}: return phrase return max(noun_phrases, key=len) if noun_phrases else None df['product_type'] = df['product_description'].apply(extract_type) print(df)
这种方法灵活性更强,但需要根据实际场景调整过滤规则,避免提取无关内容。
三、监督学习法(适合有大量标注数据的场景)
如果有足够多的标注数据(如示例中的对应关系),可以训练分类模型来自动识别商品类型:
- 用
scikit-learn构建文本分类管道,或者用BERT等深度学习模型做序列标注 - 代码示例(基于scikit-learn):
import pandas as pd from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.svm import SVC from sklearn.pipeline import Pipeline # 标注好的训练数据 train_data = { 'product_description': [ 'kanchivaram saree of red colour', 'Pink gujrati saree', 'Lehenga from Surat', 'Red swim suit', 'Blue cotton saree', 'Designer lehenga' ], 'product_type': ['saree', 'saree', 'lehenga', 'swim suit', 'saree', 'lehenga'] } train_df = pd.DataFrame(train_data) # 构建TF-IDF+SVM分类管道 pipeline = Pipeline([ ('tfidf', TfidfVectorizer(stop_words='english')), ('classifier', SVC(kernel='linear')) ]) # 训练模型 pipeline.fit(train_df['product_description'], train_df['product_type']) # 预测新数据 test_data = {'product_description': ['Green silk saree', 'Black swim suit']} test_df = pd.DataFrame(test_data) test_df['product_type'] = pipeline.predict(test_df['product_description']) print(test_df)
这种方法准确率更高,能处理复杂的描述文本,但需要足够的标注数据支撑模型训练。
内容的提问来源于stack exchange,提问作者Aditya
相关产品推荐
相关产品推荐

