基于多列修改列值及Pandas关键词匹配结果返回技术问询
Solution: Return Matching Keyword Instead of Boolean Columns
Got it, let's adjust your code so it returns the actual matching keyword (or 'blank' if no match) in a single column, instead of creating multiple boolean columns. Here's how to fix this up:
Step-by-Step Breakdown
- Replace loop with row-wise logic: Instead of generating a new boolean column for every keyword, we'll use a custom function to check each row for matches.
- Case-insensitive matching: You already converted the
Skillcolumn to uppercase, which is perfect for ensuring matches regardless of how the original text is capitalized. - Return first matching keyword: For each row, we'll scan your keyword list and return the first match we find. If no keywords are present, we'll return
'blank'.
Modified Full Code
import pandas as pd data = pd.read_excel("C:/Users/606736.CTS/Desktop/Keyword.xlsx") # Drop null value rows to avoid errors data.dropna(inplace=True) # Convert Skill column to uppercase for consistent matching data["Uppercase"] = data["Skill"].str.upper() # Your target keywords list sub = ['MEMORY', 'PASSWORD', 'DISK', 'LOGIN', 'RESET'] # Custom function to find the first matching keyword in a row def get_matching_keyword(row): row_text = row["Uppercase"] for keyword in sub: if keyword in row_text: return keyword return 'blank' # Apply the function to create a new results column data["Matching_Keyword"] = data.apply(get_matching_keyword, axis=1) # Optional: Clean up to keep only relevant columns (e.g., original Skill and the match) # data = data[["Skill", "Matching_Keyword"]] print(data.head())
Optional: Return All Matching Keywords
If you want to capture every keyword that matches a row (instead of just the first), tweak the function like this:
def get_all_matching_keywords(row): row_text = row["Uppercase"] matches = [kw for kw in sub if kw in row_text] return ', '.join(matches) if matches else 'blank' data["All_Matching_Keywords"] = data.apply(get_all_matching_keywords, axis=1)
Why This Works
- Using
apply()withaxis=1runs our custom logic across each row, keeping your DataFrame clean and focused on the result you need. - The uppercase conversion ensures matches work even if the original
Skilltext uses lowercase or mixed case (e.g., "memory" or "Memory" will still trigger a match for "MEMORY"). - This approach avoids cluttering your data with unnecessary boolean columns, giving you a direct, human-readable result.
内容的提问来源于stack exchange,提问作者Ashiq
相关产品推荐
相关产品推荐

