You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:检测文本DataFrame中客户行是否包含其他客户同义字符串的实现方案

Let's break this down step by step—you don't need multiple DataFrames, and we can even avoid clunky explicit loops (though a simple nested loop works great for small datasets) by leveraging pandas' vectorized operations and a fast lookup dictionary for synonyms. Here's how to tackle your problem:

Step 1: Build a Synonym-to-Customer Lookup Dictionary

First, convert your customer synonym data into a dictionary that maps every synonym to its parent customer. This gives you O(1) lookup speed, which is way more efficient than scanning the DataFrame every time you need to check a word:

import pandas as pd
import re

# Your original data
customer_list = {'A': ['AA', 'AA', 'AAA'], 'B': ['B', 'BB','BBB'], 'C': ['C','CC','CCC']}
customer_df = pd.DataFrame.from_dict(customer_list, orient='index')

text = [['A', 'Hello i am AA', 'Hello i am BB', 'Hello i am A'], ['B', 'Hello i am B', 'Hello i am BBB','Hello i am BB'], ['C', 'Hello i am AAA','Hello i am CC','Hello i am CCC']]
text_df = pd.DataFrame(text).set_index(0).rename_axis("customer")

# Build the lookup dict (deduplicate synonyms to avoid redundancy)
synonym_map = {}
for customer, synonyms in customer_list.items():
    for syn in set(synonyms):
        synonym_map[syn] = customer

Step 2: Create a Helper Function to Detect Errors

Write a small function that checks a single text entry for synonyms belonging to other customers. We use regex to match whole words only, so we don't get false hits (like mistaking "AA" for part of "AAA"):

def detect_error(text_entry, current_customer):
    # Match whole words only using regex
    synonym_pattern = re.compile(r'\b(' + '|'.join(re.escape(syn) for syn in synonym_map.keys()) + r')\b')
    found_synonyms = synonym_pattern.findall(text_entry)
    
    # Check if any found synonym belongs to a different customer
    for syn in found_synonyms:
        if synonym_map[syn] != current_customer:
            return f"{current_customer}: {text_entry}"
    return None

Step 3: Scan the Entire text_df for Errors

You can use pandas' applymap to run the error check on every cell in your DataFrame, then collect all valid errors (dropping None values):

# Apply the error check to every cell
error_results = text_df.applymap(lambda entry: detect_error(entry, entry.name))

# Flatten results and clean up
errors = error_results.stack().dropna().tolist()

# Print the final errors
for error in errors:
    print(f"错误文本:{error}")

Running this will output exactly what you're looking for:

错误文本:A: Hello i am BB
错误文本:C: Hello i am AAA

Alternative: Simple Nested Loop (If You Prefer Explicit Control)

If you'd rather use a loop instead of pandas' vectorized operations, this approach works just as well with only one set of nested loops (no extra DataFrames needed):

errors = []
for customer, row in text_df.iterrows():
    for text_entry in row:
        synonym_pattern = re.compile(r'\b(' + '|'.join(re.escape(syn) for syn in synonym_map.keys()) + r')\b')
        found_synonyms = synonym_pattern.findall(text_entry)
        for syn in found_synonyms:
            if synonym_map[syn] != customer:
                errors.append(f"{customer}: {text_entry}")

# Print errors
for error in errors:
    print(f"错误文本:{error}")

Key Notes for Your Implementation

  • Data Types: Use a dict for the synonym-to-customer mapping (fast lookups) and regex for precise word matching.
  • Functions: Leverage pandas' applymap for clean vectorized operations, or iterrows for explicit row-by-row processing—no need to create multiple DataFrames.
  • Loop Efficiency: You only need one set of nested loops (or no explicit loops at all with applymap) to scan all entries and detect errors.

内容的提问来源于stack exchange,提问作者DanielSchloß

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 22:03:13