You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DataFrame中匹配字符串关键词并生成对应新列标签?

Solution for Your Pandas DataFrame Tasks & Replace Error

Hey there! Let's tackle your two issues one by one—first tagging complaint categories based on keywords, then fixing that replace error.

1. Tagging Complaint Categories with Keywords

Your goal is to create a new column that marks "network issue" when the text contains "network", and "payment issue" when it contains "payment". Using replace won't work here because that's for exact string matches, not substring checks. Instead, use str.contains combined with loc to target the right rows:

import numpy as np
import pandas as pd

# Example DataFrame matching your input
data = {
    "complainttypes": [
        "Payment disappear - service got disconnected",
        "Speed and Service",
        "Provider Imposed a New Usage Cap of 300GB that ...",
        "Provider Network not working and no service to boot"
    ]
}
df = pd.DataFrame(data)

# Create new column with default "Other issue" tag
df["complaint_category"] = "Other issue"

# Tag network issues (case-insensitive match to catch "Network" or "network")
df.loc[df["complainttypes"].str.contains("network", case=False), "complaint_category"] = "network issue"

# Tag payment issues (same case-insensitive logic)
df.loc[df["complainttypes"].str.contains("payment", case=False), "complaint_category"] = "payment issue"

print(df)

This will output:

complainttypes complaint_category
0      Payment disappear - service got disconnected      payment issue
1                                Speed and Service       Other issue
2  Provider Imposed a New Usage Cap of 300GB that ...       Other issue
3  Provider Network not working and no service to boot     network issue

2. Fixing the Replace Error

The error in df["complainttypes"] = df["complainttypes"].replace({"Internet":"Internettype"}) usually stems from one of these common issues—here's how to fix each:

Common Causes & Fixes:

  • Issue 1: You're targeting substrings, not exact matches
    The default replace method looks for full-string matches. If you want to replace every occurrence of the substring "Internet", use str.replace instead:

    df["complainttypes"] = df["complainttypes"].str.replace("Internet", "Internettype", case=False)
    
  • Issue 2: Column name is misspelled or doesn't exist
    Double-check your column names with print(df.columns) to confirm "complainttypes" is spelled correctly and present.

  • Issue 3: The column contains non-string values (like NaN)
    Convert the column to string type first to avoid type errors:

    # Fill empty values and convert to string
    df["complainttypes"] = df["complainttypes"].fillna("").astype(str)
    # Run exact-match replace with regex disabled
    df["complainttypes"] = df["complainttypes"].replace({"Internet":"Internettype"}, regex=False)
    
  • Issue 4: "Internet" doesn't exist as an exact value in the column
    If you only want exact matches, confirm the value exists with print(df["complainttypes"].unique()). If it doesn't, use the substring-focused str.replace method above instead.

内容的提问来源于stack exchange,提问作者Mohammed Shaheer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 11:22:53