如何使用RASA NLU处理复合类型实体?以姓名实体场景为例
Absolutely, Rasa does support handling composite entities like the full name you mentioned (John James Doe broken into first name, middle name, and last name)! Here are a few practical approaches to implement this based on your needs:
1. Use Entity Roles (Built-in, Simplest Approach)
Rasa's native entity roles let you tag sub-parts of a composite entity with specific roles, keeping them under a parent entity type. This is perfect for your name-splitting use case.
Example Training Data
{ "rasa_nlu_data": { "common_examples": [ { "text": "My name is John James Doe", "intent": "Introduction", "entities": [ { "start": 11, "end": 15, "value": "John", "entity": "Name", "role": "first_name" }, { "start": 16, "end": 21, "value": "James", "entity": "Name", "role": "middle_name" }, { "start": 22, "end": 25, "value": "Doe", "entity": "Name", "role": "last_name" } ] } ] } }
How to Access the Data
In your custom actions or dialogue management logic, you can retrieve specific parts using the role:
# Get the first name first_name = next((e["value"] for e in tracker.latest_message["entities"] if e["entity"] == "Name" and e["role"] == "first_name"), None)
2. Build a Custom NER Component (For Automatic Splitting)
If you don't want to manually tag every sub-part of the composite entity, you can create a custom component that splits a recognized parent entity (like the full Name entity from your original data) into its sub-components automatically.
Step 1: Create the Custom Component
from rasa.nlu.components import Component from rasa.shared.nlu.training_data.message import Message from typing import List, Optional, Dict class SplitFullNameComponent(Component): def __init__(self, component_config: Optional[Dict] = None): super().__init__(component_config) def process(self, messages: List[Message], **kwargs): for message in messages: entities = message.get("entities", []) updated_entities = [] for entity in entities: if entity["entity"] == "Name": name_parts = entity["value"].split() name_length = len(name_parts) # Add first name first_start = entity["start"] first_end = first_start + len(name_parts[0]) updated_entities.append({ "start": first_start, "end": first_end, "value": name_parts[0], "entity": "first_name", "extractor": "custom_name_splitter" }) # Add middle names (if any) if name_length > 2: current_pos = first_end + 1 for part in name_parts[1:-1]: mid_start = current_pos mid_end = mid_start + len(part) updated_entities.append({ "start": mid_start, "end": mid_end, "value": part, "entity": "middle_name", "extractor": "custom_name_splitter" }) current_pos = mid_end + 1 # Add last name last_start = entity["end"] - len(name_parts[-1]) updated_entities.append({ "start": last_start, "end": entity["end"], "value": name_parts[-1], "entity": "last_name", "extractor": "custom_name_splitter" }) else: updated_entities.append(entity) message.set("entities", updated_entities, add_to_output=True)
Step 2: Register the Component in config.yml
Add your custom component to the NLU pipeline after the standard NER components:
language: en pipeline: - name: WhitespaceTokenizer - name: LexicalSyntacticFeaturizer - name: CountVectorsFeaturizer - name: DIETClassifier epochs: 100 - name: EntitySynonymMapper - name: "path.to.your.SplitFullNameComponent" # Replace with your actual file path
3. Use Entity Groups (For Associating Sub-Entities)
If you need to link multiple sub-entities to indicate they belong to the same composite entity (e.g., distinguishing between two people's names in one message), use entity groups.
Example Training Data
{ "rasa_nlu_data": { "common_examples": [ { "text": "My name is John James Doe", "intent": "Introduction", "entities": [ { "start": 11, "end": 15, "value": "John", "entity": "first_name", "group": "full_name_1" }, { "start": 16, "end": 21, "value": "James", "entity": "middle_name", "group": "full_name_1" }, { "start": 22, "end": 25, "value": "Doe", "entity": "last_name", "group": "full_name_1" } ] } ] } }
This lets you group all sub-entities under a single identifier, making it easy to associate them in your logic.
内容的提问来源于stack exchange,提问作者Kunal Mukherjee

