如何在DataFrame中匹配字符串关键词并生成对应新列标签?
Hey there! Let's tackle your two issues one by one—first tagging complaint categories based on keywords, then fixing that replace error.
1. Tagging Complaint Categories with Keywords
Your goal is to create a new column that marks "network issue" when the text contains "network", and "payment issue" when it contains "payment". Using replace won't work here because that's for exact string matches, not substring checks. Instead, use str.contains combined with loc to target the right rows:
import numpy as np import pandas as pd # Example DataFrame matching your input data = { "complainttypes": [ "Payment disappear - service got disconnected", "Speed and Service", "Provider Imposed a New Usage Cap of 300GB that ...", "Provider Network not working and no service to boot" ] } df = pd.DataFrame(data) # Create new column with default "Other issue" tag df["complaint_category"] = "Other issue" # Tag network issues (case-insensitive match to catch "Network" or "network") df.loc[df["complainttypes"].str.contains("network", case=False), "complaint_category"] = "network issue" # Tag payment issues (same case-insensitive logic) df.loc[df["complainttypes"].str.contains("payment", case=False), "complaint_category"] = "payment issue" print(df)
This will output:
complainttypes complaint_category 0 Payment disappear - service got disconnected payment issue 1 Speed and Service Other issue 2 Provider Imposed a New Usage Cap of 300GB that ... Other issue 3 Provider Network not working and no service to boot network issue
2. Fixing the Replace Error
The error in df["complainttypes"] = df["complainttypes"].replace({"Internet":"Internettype"}) usually stems from one of these common issues—here's how to fix each:
Common Causes & Fixes:
Issue 1: You're targeting substrings, not exact matches
The defaultreplacemethod looks for full-string matches. If you want to replace every occurrence of the substring "Internet", usestr.replaceinstead:df["complainttypes"] = df["complainttypes"].str.replace("Internet", "Internettype", case=False)Issue 2: Column name is misspelled or doesn't exist
Double-check your column names withprint(df.columns)to confirm "complainttypes" is spelled correctly and present.Issue 3: The column contains non-string values (like NaN)
Convert the column to string type first to avoid type errors:# Fill empty values and convert to string df["complainttypes"] = df["complainttypes"].fillna("").astype(str) # Run exact-match replace with regex disabled df["complainttypes"] = df["complainttypes"].replace({"Internet":"Internettype"}, regex=False)Issue 4: "Internet" doesn't exist as an exact value in the column
If you only want exact matches, confirm the value exists withprint(df["complainttypes"].unique()). If it doesn't, use the substring-focusedstr.replacemethod above instead.
内容的提问来源于stack exchange,提问作者Mohammed Shaheer

