You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何清洗DataFrame多行业列,避免无序行业组合重复统计?

多行业列拆分统计解决方案
  • 处理字符串格式的行业数组(若你的列是字符串形式的列表):
    用ast.literal_eval把字符串转成真实列表,避免后续操作无效:
    import ast
    df['cust_industry'] = df['cust_industry'].apply(ast.literal_eval)
    
  • 拆分多行业组合:
    用explode方法将每个客户的多个行业拆成单独行,每个行业对应一条记录:
    df_exploded = df.explode('cust_industry', ignore_index=True)
    
  • 标准化行业名称:
    统一大小写、去除冗余空格,解决authorized和Authorized被误判为不同类别的问题:
    df_exploded['cust_industry'] = df_exploded['cust_industry'].str.lower().str.strip()
    
    若存在同义词/不同写法,用字典映射统一:
    name_map = {
        'authorized': 'authorized',
        'distribution': 'distribution'
        # 按需添加其他需要统一的名称
    }
    df_exploded['cust_industry'] = df_exploded['cust_industry'].replace(name_map)
    
  • 按单个行业统计:
    用value_counts或groupby直接统计,结果就是每个行业的独立计数:
    # 简洁版
    industry_stats = df_exploded['cust_industry'].value_counts().reset_index(name='count')
    
    # 或者groupby版本
    industry_stats = df_exploded.groupby('cust_industry').size().reset_index(name='count')
    

内容的提问来源于stack exchange,提问作者Sob.Ruba

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 23:25:12