Python:检测文本DataFrame中客户行是否包含其他客户同义字符串的实现方案
Let's break this down step by step—you don't need multiple DataFrames, and we can even avoid clunky explicit loops (though a simple nested loop works great for small datasets) by leveraging pandas' vectorized operations and a fast lookup dictionary for synonyms. Here's how to tackle your problem:
Step 1: Build a Synonym-to-Customer Lookup Dictionary
First, convert your customer synonym data into a dictionary that maps every synonym to its parent customer. This gives you O(1) lookup speed, which is way more efficient than scanning the DataFrame every time you need to check a word:
import pandas as pd import re # Your original data customer_list = {'A': ['AA', 'AA', 'AAA'], 'B': ['B', 'BB','BBB'], 'C': ['C','CC','CCC']} customer_df = pd.DataFrame.from_dict(customer_list, orient='index') text = [['A', 'Hello i am AA', 'Hello i am BB', 'Hello i am A'], ['B', 'Hello i am B', 'Hello i am BBB','Hello i am BB'], ['C', 'Hello i am AAA','Hello i am CC','Hello i am CCC']] text_df = pd.DataFrame(text).set_index(0).rename_axis("customer") # Build the lookup dict (deduplicate synonyms to avoid redundancy) synonym_map = {} for customer, synonyms in customer_list.items(): for syn in set(synonyms): synonym_map[syn] = customer
Step 2: Create a Helper Function to Detect Errors
Write a small function that checks a single text entry for synonyms belonging to other customers. We use regex to match whole words only, so we don't get false hits (like mistaking "AA" for part of "AAA"):
def detect_error(text_entry, current_customer): # Match whole words only using regex synonym_pattern = re.compile(r'\b(' + '|'.join(re.escape(syn) for syn in synonym_map.keys()) + r')\b') found_synonyms = synonym_pattern.findall(text_entry) # Check if any found synonym belongs to a different customer for syn in found_synonyms: if synonym_map[syn] != current_customer: return f"{current_customer}: {text_entry}" return None
Step 3: Scan the Entire text_df for Errors
You can use pandas' applymap to run the error check on every cell in your DataFrame, then collect all valid errors (dropping None values):
# Apply the error check to every cell error_results = text_df.applymap(lambda entry: detect_error(entry, entry.name)) # Flatten results and clean up errors = error_results.stack().dropna().tolist() # Print the final errors for error in errors: print(f"错误文本:{error}")
Running this will output exactly what you're looking for:
错误文本:A: Hello i am BB 错误文本:C: Hello i am AAA
Alternative: Simple Nested Loop (If You Prefer Explicit Control)
If you'd rather use a loop instead of pandas' vectorized operations, this approach works just as well with only one set of nested loops (no extra DataFrames needed):
errors = [] for customer, row in text_df.iterrows(): for text_entry in row: synonym_pattern = re.compile(r'\b(' + '|'.join(re.escape(syn) for syn in synonym_map.keys()) + r')\b') found_synonyms = synonym_pattern.findall(text_entry) for syn in found_synonyms: if synonym_map[syn] != customer: errors.append(f"{customer}: {text_entry}") # Print errors for error in errors: print(f"错误文本:{error}")
Key Notes for Your Implementation
- Data Types: Use a
dictfor the synonym-to-customer mapping (fast lookups) and regex for precise word matching. - Functions: Leverage pandas'
applymapfor clean vectorized operations, oriterrowsfor explicit row-by-row processing—no need to create multiple DataFrames. - Loop Efficiency: You only need one set of nested loops (or no explicit loops at all with
applymap) to scan all entries and detect errors.
内容的提问来源于stack exchange,提问作者DanielSchloß

