You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Brat标注配置设计内部工具约束定义语言(Python实现)

Designing a Brat-like Constraint Validation Language with a Python Parser

Understanding the Core Config Structure

First, let's break down the Brat-style config you provided to map out exactly what we need to support:

Entities Section

The [entities] block defines all valid entity types your tool will recognize. For your example:

[entities] Drug DrugClass Procedure Therapy AE SAE Disease

This is a simple whitespace-separated list of allowed entity names.

Relations Section

The [relations] block specifies valid relationships between entities, plus constraints:

[relations]
Equiv Arg1:, Arg2:, :symmetric-transitive
BelongsTo Arg1:Drug , Arg2:DrugClass
BelongsTo Arg1:AE , ...

Each relation entry includes:

  • A unique relation name (e.g., Equiv, BelongsTo)
  • Argument constraints (e.g., Arg1:Drug means the first argument must be a Drug; <ENTITY> accepts any valid entity)
  • Optional relation properties (e.g., symmetric-transitive enforces those behaviors for Equiv)

Step 1: Define a Data Model

We'll start with Python dataclasses to represent parsed config elements—this makes working with the data clean and intuitive.

from dataclasses import dataclass
from typing import Set, Dict, List, Optional

@dataclass
class RelationArg:
    arg_id: str  # e.g., "Arg1"
    required_entity: str  # e.g., "Drug" or "<ENTITY>"

@dataclass
class RelationRule:
    name: str  # e.g., "Equiv"
    args: List[RelationArg]
    properties: Optional[Set[str]] = None  # e.g., {"symmetric", "transitive"}

@dataclass
class AnnotationConfig:
    entities: Set[str]
    relations: Dict[str, RelationRule]

Step 2: Build the Parser

Next, write a parser that converts the raw config string into our data model. We'll split the config into sections, then process each line to extract entities and relations.

def parse_config(config_text: str) -> AnnotationConfig:
    lines = [line.strip() for line in config_text.splitlines() if line.strip()]
    current_section = None
    entities = set()
    relations = {}

    for line in lines:
        # Switch sections when we hit [section] tags
        if line.startswith("[") and line.endswith("]"):
            current_section = line.strip("[]").lower()
            continue
        
        if current_section == "entities":
            # Add all whitespace-separated entities to the set
            entities.update(line.split())
        
        elif current_section == "relations":
            # Split line into parts, ignoring extra spaces around commas
            parts = [p.strip() for p in line.split(",") if p.strip()]
            # First part has the relation name plus initial args
            first_part = parts[0].split()
            rel_name = first_part[0]
            rel_components = first_part[1:] + parts[1:]

            args = []
            props = set()

            for component in rel_components:
                key, value = component.split(":", 1)
                key = key.strip()
                value = value.strip()

                if key.startswith("Arg"):
                    args.append(RelationArg(arg_id=key, required_entity=value))
                elif key == "<REL-TYPE>":
                    # Split hyphenated properties into individual flags
                    props.update(value.split("-"))
            
            relations[rel_name] = RelationRule(
                name=rel_name,
                args=args,
                properties=props if props else None
            )
    
    return AnnotationConfig(entities=entities, relations=relations)

Step 3: Add Validation Logic

Now use the parsed config to validate annotations. This example checks if a relation is allowed, if arguments match required entity types, and enforces basic property rules.

def validate_relation(config: AnnotationConfig, rel_name: str, arg1_type: str, arg2_type: str) -> bool:
    # Check if the relation exists in the config
    if rel_name not in config.relations:
        print(f"Invalid relation: '{rel_name}' is not defined")
        return False
    
    rel_rule = config.relations[rel_name]

    # Validate argument count (assuming 2-arg relations for simplicity)
    if len(rel_rule.args) != 2:
        print(f"Relation '{rel_name}' requires {len(rel_rule.args)} arguments")
        return False
    
    # Validate Arg1 entity type
    arg1_req = rel_rule.args[0].required_entity
    if arg1_req != "<ENTITY>" and arg1_type != arg1_req:
        print(f"Arg1 for '{rel_name}' must be '{arg1_req}', got '{arg1_type}'")
        return False
    
    # Validate Arg2 entity type
    arg2_req = rel_rule.args[1].required_entity
    if arg2_req != "<ENTITY>" and arg2_type != arg2_req:
        print(f"Arg2 for '{rel_name}' must be '{arg2_req}', got '{arg2_type}'")
        return False
    
    # Check if entities are valid
    if arg1_type not in config.entities:
        print(f"Invalid entity: '{arg1_type}' is not defined")
        return False
    if arg2_type not in config.entities:
        print(f"Invalid entity: '{arg2_type}' is not defined")
        return False
    
    # Enforce symmetry if specified
    if "symmetric" in rel_rule.properties and not validate_relation(config, rel_name, arg2_type, arg1_type):
        print(f"Relation '{rel_name}' must be symmetric")
        return False
    
    return True

# Example usage
demo_config = """
[entities] Drug DrugClass Procedure Therapy AE SAE Disease
[relations] 
Equiv Arg1:<ENTITY>, Arg2:<ENTITY>, <REL-TYPE>:symmetric-transitive 
BelongsTo Arg1:Drug , Arg2:DrugClass 
BelongsTo Arg1:AE , Arg2:Disease
"""

parsed = parse_config(demo_config)
# Valid case
print(validate_relation(parsed, "BelongsTo", "Drug", "DrugClass"))  # Output: True
# Invalid case (wrong Arg1 type)
print(validate_relation(parsed, "BelongsTo", "DrugClass", "Drug"))  # Output: False

Step 4: Extend for Your Needs

You can expand this foundation with:

  • Support for multi-argument relations
  • More complex entity constraints (e.g., wildcards, subsets)
  • Detailed error reporting with line numbers
  • Serialization to/from JSON/YAML for easier storage
  • Integration with your internal tool's annotation UI

内容的提问来源于stack exchange,提问作者vanangamudi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:26:58