You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas批量处理CSV时,新增列因条件逻辑错误致值统一问题

问题解决:Pandas批量处理CSV时新增Device Type列的错误修复

问题描述

批量处理CSV文件时,尝试基于Product Type列包含的特定字符串新增Device Type列,但出现异常:整列Device Type会被统一设为首个满足条件的值,比如首行是HW-VG54-NAH,整列都变成Gateway,完全忽略其他行的不同产品类型。

错误原因

原代码中conditions里误用了.any()方法:

conditions = [
    (df['Product Type'].str.contains('VG').any()),
    (df['Product Type'].str.contains('AG').any()),
    (df['Product Type'].str.contains('CM3').any())
]

.any()会对整个列做全局判断,返回单一布尔值(只要列中有一个元素满足条件就返回True),导致np.select将整列都赋值为第一个匹配的value,而非逐行判断每行的产品类型。

修复后的完整代码

import glob
import os
import pandas as pd
import numpy as np

path = "你的CSV文件路径"  # 替换为实际路径
csv_files = glob.glob(os.path.join(path, "*.csv"))
print('Found', len(csv_files), 'files')

# 需要校验的表头
header_list = ['Created Date', 'Order Number', 'Shipping Address', 'Shipping Contact', 'Shipping Email', 'Product Type', 'Quantity', 'Serial', 'Activation Status']

list_of_df = []

for f in csv_files: 
    print('File Name:', f.split("\\")[-1])
    
    # 读取文件
    df = pd.read_csv(f, index_col=None, header=0)

    # 修改条件:移除.any(),实现逐行判断
    conditions = [
        df['Product Type'].str.contains('VG', na=False),
        df['Product Type'].str.contains('AG', na=False),
        df['Product Type'].str.contains('CM3', na=False)
    ]
    values = ['Gateway', 'Asset', 'Camera']

    # 校验表头
    import_headers = df.columns
    missing_headers = [i for i in import_headers if i not in header_list]
    
    if not missing_headers:
        print('Headers are OK, file is imported')
        # 数据预处理
        # 删除指定列
        df.drop(df.columns[[2,3,4]], axis=1, inplace=True)

        # 填充并转换Activation Status列
        df["Activation Status"] = df["Activation Status"].fillna("0")
        df['Activation Status'] = df['Activation Status'].replace('Activated', '1')
        
        # 筛选包含HW-的行
        df = df.loc[df['Product Type'].str.contains('HW-', regex=True, na=False)]
        
        # 新增Device Type列(逐行匹配)
        df['Device Type'] = np.select(conditions, values, default='Unknown')  # 可选:添加默认值处理未匹配项
        
        list_of_df.append(df)
    else:
        print('Headers are not OK, file is not imported')
        print('Headers not found:', missing_headers)
        print('Headers found:', import_headers)

# 合并所有处理后的数据集
final_df = pd.concat(list_of_df, axis=0, ignore_index=True)

关键修改说明

  1. 移除.any():让str.contains()返回逐行的布尔值Series,确保np.select能为每行匹配对应的Device Type。
  2. 添加na=False:避免空值导致的判断异常,确保空值不会被误判为满足条件。
  3. 可选默认值:通过default参数为不匹配任何条件的行设置默认值(如'Unknown'),避免出现NaN。

内容的提问来源于stack exchange,提问作者AdrianM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 11:23:30