You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas交易数据按供应商匹配分类列赋值问题求助

How to Fix Pandas Category Assignment for Transaction Data

Let's break down why your Category column is ending up all NaN and fix it step by step.

What's Wrong with Your Current Code

  1. Backwards Code Order: You’re trying to loop through categories_dict before you even define it—this should throw an error right off the bat. Even if you rearranged it accidentally, the bigger issue is:
  2. Function Overwriting in Loops: You’re defining the to_key function inside a nested loop. Every iteration replaces to_key with a new version that only checks the current vendor and returns the current category. By the time you run apply(), to_key only looks for the last vendor from the last category (VENDOR9), which isn’t in your data—hence all NaNs.

Fixed Solutions

Solution 1: Dedicated Function for Full Vendor Check

First, define your category dictionary before using it, then write a function that scans each description against all vendors:

# Define the category dictionary FIRST
categories_dict = {
    'category1': ['VENDOR1', 'VENDOR2', 'VENDOR3'],
    'category2': ['VENDOR4', 'VENDOR5', 'VENDOR6'],
    'category3': ['VENDOR7', 'VENDOR8', 'VENDOR9']
}

# Function to map description to category
def get_transaction_category(description):
    for category, vendors in categories_dict.items():
        for vendor in vendors:
            if vendor in description:
                return category
    # Return None if no match (becomes NaN in pandas)
    return None

# Apply the function to the Description column
df["Category"] = df["Description"].apply(get_transaction_category)

Solution 2: Optimized Vendor-to-Category Mapping

For larger datasets, building a reverse mapping first speeds up lookups:

categories_dict = {
    'category1': ['VENDOR1', 'VENDOR2', 'VENDOR3'],
    'category2': ['VENDOR4', 'VENDOR5', 'VENDOR6'],
    'category3': ['VENDOR7', 'VENDOR8', 'VENDOR9']
}

# Build a reverse dictionary: vendor -> category
vendor_category_map = {}
for cat, vendors in categories_dict.items():
    for vendor in vendors:
        vendor_category_map[vendor] = cat

def get_transaction_category(description):
    for vendor, category in vendor_category_map.items():
        if vendor in description:
            return category
    return None

df["Category"] = df["Description"].apply(get_transaction_category)

Solution 3: Concise Version with next()

For a more compact approach, use a generator expression with next() to find the first matching category:

categories_dict = {
    'category1': ['VENDOR1', 'VENDOR2', 'VENDOR3'],
    'category2': ['VENDOR4', 'VENDOR5', 'VENDOR6'],
    'category3': ['VENDOR7', 'VENDOR8', 'VENDOR9']
}

df["Category"] = df["Description"].apply(
    lambda desc: next(
        (cat for cat, vendors in categories_dict.items() 
         for v in vendors if v in desc),
        None
    )
)

Final Result

After running any of these solutions, your DataFrame will match your expected output:

DateDescriptionAmountCategory
01/11/20VENDOR1 #34299.54category1
05/11/20VENDOR2 #762100.5category1
06/11/20VENDOR4 #32116.54category2
06/11/20VENDOR12 #5732.54NaN
09/11/20VENDOR7 #22275.54category3

内容的提问来源于stack exchange,提问作者Christopher Vaux

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:21:52