You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python提取法语CSV商品数据中的纯颜色名称

提取法语CSV中的纯颜色名称并统计热销颜色

核心思路

针对法语颜色字段的特点,我们可以先建立一个法语常见颜色词库,再通过匹配词库的方式从带描述的字段中提取纯颜色名称,最后统计各颜色的出现频次来确定热销颜色。

实现步骤及代码

  1. 定义法语颜色词库:覆盖常用基础颜色及常见变体,可根据你的数据扩展
  2. 读取CSV文件并处理每个颜色字段
  3. 统计各颜色的出现次数
import pandas as pd

# 定义法语常用颜色集合,可根据实际数据补充
french_colors = {
    "Bleu", "Marron", "Rouge", "Vert", "Jaune", "Noir", "Blanc", 
    "Gris", "Orange", "Rose", "Violet", "Turquoise", "Beige", "Doré", "Argenté",
    "Bleu clair", "Bleu foncé", "Rouge clair", "Vert clair"
}

def extract_pure_color(color_text):
    # 处理空值
    if pd.isna(color_text):
        return None
    # 先尝试匹配复合颜色(比如带修饰的颜色)
    for color in french_colors:
        if color in color_text:
            return color
    # 如果没有复合颜色,拆分单词匹配基础颜色
    words = [word.strip() for word in color_text.split()]
    for word in words:
        if word in french_colors:
            return word
    # 未匹配到颜色时返回原文本或None,按需调整
    return None

# 读取CSV文件,替换为你的文件路径和颜色列名
df = pd.read_csv("products.csv")
# 假设颜色列名为"couleur"(法语常用列名),替换为你的实际列名
df["pure_couleur"] = df["couleur"].apply(extract_pure_color)

# 统计热销颜色,按出现次数降序排列
hot_colors = df["pure_couleur"].value_counts().dropna()
print("热销颜色统计:")
print(hot_colors)

细节调整建议

  • 如果你的数据里有带连字符的颜色(比如Bleu-clair),可以在处理前先把连字符替换为空格:color_text = color_text.replace("-", " ")
  • 如果存在一个字段包含多个颜色的情况(比如Bleu et Rouge),可以修改函数返回匹配到的所有颜色,再用df.explode()拆分后统计:
    def extract_multiple_colors(color_text):
        if pd.isna(color_text):
            return []
        color_text = color_text.replace("-", " ")
        matched = []
        # 先匹配复合颜色
        for color in [c for c in french_colors if len(c.split())>1]:
            if color in color_text:
                matched.append(color)
                # 避免重复匹配基础颜色
                color_text = color_text.replace(color, "")
        # 匹配基础颜色
        words = [word.strip() for word in color_text.split()]
        matched.extend([w for w in words if w in french_colors])
        return matched
    
    df["pure_couleur"] = df["couleur"].apply(extract_multiple_colors)
    hot_colors = df.explode("pure_couleur")["pure_couleur"].value_counts().dropna()
    
  • 可以根据你的实际数据不断补充french_colors集合,提升匹配准确率

内容的提问来源于stack exchange,提问作者Marie A

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 10:25:19