如何清洗DataFrame多行业列,避免无序行业组合重复统计?
多行业列拆分统计解决方案
- 处理字符串格式的行业数组(若你的列是字符串形式的列表):
用ast.literal_eval把字符串转成真实列表,避免后续操作无效:import ast df['cust_industry'] = df['cust_industry'].apply(ast.literal_eval) - 拆分多行业组合:
用explode方法将每个客户的多个行业拆成单独行,每个行业对应一条记录:df_exploded = df.explode('cust_industry', ignore_index=True) - 标准化行业名称:
统一大小写、去除冗余空格,解决authorized和Authorized被误判为不同类别的问题:
若存在同义词/不同写法,用字典映射统一:df_exploded['cust_industry'] = df_exploded['cust_industry'].str.lower().str.strip()name_map = { 'authorized': 'authorized', 'distribution': 'distribution' # 按需添加其他需要统一的名称 } df_exploded['cust_industry'] = df_exploded['cust_industry'].replace(name_map) - 按单个行业统计:
用value_counts或groupby直接统计,结果就是每个行业的独立计数:# 简洁版 industry_stats = df_exploded['cust_industry'].value_counts().reset_index(name='count') # 或者groupby版本 industry_stats = df_exploded.groupby('cust_industry').size().reset_index(name='count')
内容的提问来源于stack exchange,提问作者Sob.Ruba
相关产品推荐
相关产品推荐

