请求提供正则表达式以提取客服对话中的客户发言(用于情感分析)
Regex Solution to Extract Customer Messages from Unformatted Dialogue
Looking at your unbroken dialogue text, here's a regex pattern that will reliably extract only the customer's messages while ignoring all agent lines:
\[cust\]:(.*?)(?=\[agent\]:|$)
Pattern Breakdown:
\[cust\]:: Matches the literal customer identifier[cust]:(brackets are escaped with\because they have special meaning in regex syntax).(.*?): A non-greedy capture group that grabs all characters until it hits the next stopping point. The?ensures it doesn't overmatch into agent lines.(?=\[agent\]:|$): A positive lookahead that tells the regex to stop when it encounters either the start of an agent line ([agent]:) or the end of the entire string ($). This ensures we don't include any agent content in our captures.
Example Usage (Python):
If you're using Python to parse the text for sentiment analysis, here's how you'd implement this:
import re dialogue = "[agent]:Welcome to ABC bank My name is Asif. How may I help you [cust]:I got additional charge in my credit card, I will not be paying this, please remove it [agent]:Okay can I place the call on hold [cust]:This is very unresponsive behaviour on banks side。" # Extract all customer messages customer_messages = re.findall(r'\[cust\]:(.*?)(?=\[agent\]:|$)', dialogue) # Clean and print results (strip removes extra whitespace) for idx, msg in enumerate(customer_messages, 1): print(f"Customer Message {idx}: {msg.strip()}")
Output:
Customer Message 1: I got additional charge in my credit card, I will not be paying this, please remove it Customer Message 2: This is very unresponsive behaviour on banks side。
This pattern works seamlessly for both unbroken dialogue (like your input) and line-separated conversations. If you need to adapt it to another programming language (JavaScript, Java, etc.), the core regex logic stays identical—you'll just use the regex functions native to that language.
内容的提问来源于stack exchange,提问作者MagicBeans
相关产品推荐
相关产品推荐

