如何用Python基于关键词为表格列文本分类?
Python Solution to Categorize Text with Keyword Rules
Here's a straightforward, pandas-based approach to replicate that Excel keyword categorization logic in Python—perfect for processing your table data:
Step 1: Define Your Keyword-Category Mapping
First, set up a dictionary that links each category to its associated keywords. You can easily expand this as your rules grow:
category_keyword_map = { "AAAA": ["keyword1", "keyword2", "keyword3"], "BBBB": ["keyword4"] }
Step 2: Create a Categorization Helper Function
This function checks a single description against your keyword rules and returns the matching category. It handles case insensitivity by default (adjustable if needed):
def get_category(description, keyword_map): # Normalize text to lowercase for consistent matching desc_lower = description.lower() # Check each category in order (priority goes to the first matching category) for category, keywords in keyword_map.items(): if any(keyword.lower() in desc_lower for keyword in keywords): return category # Fallback if no keywords are found return "Uncategorized"
Step 3: Apply to Your DataFrame
Use pandas to load your data (from Excel, CSV, etc.) and apply the function to generate the category column:
import pandas as pd # Sample input data (replace with your actual data source) data = { "Description": [ "this free text keyword1 is done", "free sample2 keyword4 keyword3", "random text with no matching keywords" ] } df = pd.DataFrame(data) # Add the category column df["category"] = df["Description"].apply(lambda x: get_category(x, category_keyword_map)) # View the result print(df)
Output:
Description category 0 this free text keyword1 is done AAAA 1 free sample2 keyword4 keyword3 BBBB 2 random text with no matching keywords Uncategorized
Customization Tips:
- Exact Word Matching: If you want to match whole words only (not substrings), modify the check to use regex word boundaries:
import re if any(re.search(rf"\b{re.escape(keyword.lower())}\b", desc_lower) for keyword in keywords): - Adjust Priority: The function returns the first category that has a matching keyword. Reorder the dictionary entries to change which category takes precedence when multiple keywords overlap.
- Load/Save Excel Files: Use
pd.read_excel("your_file.xlsx")to load your data anddf.to_excel("categorized_file.xlsx", index=False)to save the result back to Excel.
内容的提问来源于stack exchange,提问作者Mouad
相关产品推荐
相关产品推荐

