基于Brat标注配置设计内部工具约束定义语言(Python实现)
Understanding the Core Config Structure
First, let's break down the Brat-style config you provided to map out exactly what we need to support:
Entities Section
The [entities] block defines all valid entity types your tool will recognize. For your example:
[entities] Drug DrugClass Procedure Therapy AE SAE Disease
This is a simple whitespace-separated list of allowed entity names.
Relations Section
The [relations] block specifies valid relationships between entities, plus constraints:
[relations]
Equiv Arg1:, Arg2: , :symmetric-transitive
BelongsTo Arg1:Drug , Arg2:DrugClass
BelongsTo Arg1:AE , ...
Each relation entry includes:
- A unique relation name (e.g.,
Equiv,BelongsTo) - Argument constraints (e.g.,
Arg1:Drugmeans the first argument must be aDrug;<ENTITY>accepts any valid entity) - Optional relation properties (e.g.,
symmetric-transitiveenforces those behaviors forEquiv)
Step 1: Define a Data Model
We'll start with Python dataclasses to represent parsed config elements—this makes working with the data clean and intuitive.
from dataclasses import dataclass from typing import Set, Dict, List, Optional @dataclass class RelationArg: arg_id: str # e.g., "Arg1" required_entity: str # e.g., "Drug" or "<ENTITY>" @dataclass class RelationRule: name: str # e.g., "Equiv" args: List[RelationArg] properties: Optional[Set[str]] = None # e.g., {"symmetric", "transitive"} @dataclass class AnnotationConfig: entities: Set[str] relations: Dict[str, RelationRule]
Step 2: Build the Parser
Next, write a parser that converts the raw config string into our data model. We'll split the config into sections, then process each line to extract entities and relations.
def parse_config(config_text: str) -> AnnotationConfig: lines = [line.strip() for line in config_text.splitlines() if line.strip()] current_section = None entities = set() relations = {} for line in lines: # Switch sections when we hit [section] tags if line.startswith("[") and line.endswith("]"): current_section = line.strip("[]").lower() continue if current_section == "entities": # Add all whitespace-separated entities to the set entities.update(line.split()) elif current_section == "relations": # Split line into parts, ignoring extra spaces around commas parts = [p.strip() for p in line.split(",") if p.strip()] # First part has the relation name plus initial args first_part = parts[0].split() rel_name = first_part[0] rel_components = first_part[1:] + parts[1:] args = [] props = set() for component in rel_components: key, value = component.split(":", 1) key = key.strip() value = value.strip() if key.startswith("Arg"): args.append(RelationArg(arg_id=key, required_entity=value)) elif key == "<REL-TYPE>": # Split hyphenated properties into individual flags props.update(value.split("-")) relations[rel_name] = RelationRule( name=rel_name, args=args, properties=props if props else None ) return AnnotationConfig(entities=entities, relations=relations)
Step 3: Add Validation Logic
Now use the parsed config to validate annotations. This example checks if a relation is allowed, if arguments match required entity types, and enforces basic property rules.
def validate_relation(config: AnnotationConfig, rel_name: str, arg1_type: str, arg2_type: str) -> bool: # Check if the relation exists in the config if rel_name not in config.relations: print(f"Invalid relation: '{rel_name}' is not defined") return False rel_rule = config.relations[rel_name] # Validate argument count (assuming 2-arg relations for simplicity) if len(rel_rule.args) != 2: print(f"Relation '{rel_name}' requires {len(rel_rule.args)} arguments") return False # Validate Arg1 entity type arg1_req = rel_rule.args[0].required_entity if arg1_req != "<ENTITY>" and arg1_type != arg1_req: print(f"Arg1 for '{rel_name}' must be '{arg1_req}', got '{arg1_type}'") return False # Validate Arg2 entity type arg2_req = rel_rule.args[1].required_entity if arg2_req != "<ENTITY>" and arg2_type != arg2_req: print(f"Arg2 for '{rel_name}' must be '{arg2_req}', got '{arg2_type}'") return False # Check if entities are valid if arg1_type not in config.entities: print(f"Invalid entity: '{arg1_type}' is not defined") return False if arg2_type not in config.entities: print(f"Invalid entity: '{arg2_type}' is not defined") return False # Enforce symmetry if specified if "symmetric" in rel_rule.properties and not validate_relation(config, rel_name, arg2_type, arg1_type): print(f"Relation '{rel_name}' must be symmetric") return False return True # Example usage demo_config = """ [entities] Drug DrugClass Procedure Therapy AE SAE Disease [relations] Equiv Arg1:<ENTITY>, Arg2:<ENTITY>, <REL-TYPE>:symmetric-transitive BelongsTo Arg1:Drug , Arg2:DrugClass BelongsTo Arg1:AE , Arg2:Disease """ parsed = parse_config(demo_config) # Valid case print(validate_relation(parsed, "BelongsTo", "Drug", "DrugClass")) # Output: True # Invalid case (wrong Arg1 type) print(validate_relation(parsed, "BelongsTo", "DrugClass", "Drug")) # Output: False
Step 4: Extend for Your Needs
You can expand this foundation with:
- Support for multi-argument relations
- More complex entity constraints (e.g., wildcards, subsets)
- Detailed error reporting with line numbers
- Serialization to/from JSON/YAML for easier storage
- Integration with your internal tool's annotation UI
内容的提问来源于stack exchange,提问作者vanangamudi

