XML标注数据集转换为BRAT格式的技术实现咨询
Got it, let's walk through how to convert your XML-annotated text to BRAT's standoff format properly. I've dealt with this exact task for medical NLP datasets before, so here's a practical, tested solution:
Step 1: Remember BRAT's Core Requirements
First, BRAT needs two paired files for every document:
- A plain text file (
.txt) with all XML tags stripped out (this is the base text BRAT displays) - An annotation file (
.ann) with line-by-line standoff entries, following the format:T<ID> <TYPE> <START_OFFSET> <END_OFFSET> <ANNOTATED_TEXT>
The tricky part is calculating accurate character offsets—since removing XML tags shifts the position of your annotated text, you can't just use positions from the original XML.
Step 2: Python Script to Handle Conversion
This script uses regex to locate annotations and tracks character positions as it builds the tag-free text, ensuring offsets are 100% accurate:
import re def xml_to_brat(xml_input, txt_out_path, ann_out_path): # Regex to match annotation tags and capture CUI + annotated text annotation_pattern = re.compile(r'<annotation cui="(.*?)">(.*?)</annotation>') plain_text_parts = [] brat_annotations = [] current_text_length = 0 annotation_id = 1 # Iterate through every matched annotation in the XML for match in annotation_pattern.finditer(xml_input): cui, annotated_phrase = match.groups() # Add the text before the current annotation to our plain text plain_text_parts.append(xml_input[current_text_length:match.start()]) # Calculate where the annotated phrase starts in the final plain text start_offset = len(''.join(plain_text_parts)) # Add the annotated phrase to the plain text plain_text_parts.append(annotated_phrase) # Calculate the end offset end_offset = start_offset + len(annotated_phrase) # Build the BRAT annotation line (use CUI as type if preferred) brat_annotations.append(f"T{annotation_id} annotation {start_offset} {end_offset} {annotated_phrase}") annotation_id += 1 # Update our position in the original XML text current_text_length = match.end() # Add any remaining text after the last annotation plain_text_parts.append(xml_input[current_text_length:]) final_plain_text = ''.join(plain_text_parts) # Write output files with open(txt_out_path, 'w', encoding='utf-8') as txt_file: txt_file.write(final_plain_text) with open(ann_out_path, 'w', encoding='utf-8') as ann_file: ann_file.write('\n'.join(brat_annotations)) # Test with your example input sample_xml = 'Treatment of <annotation cui="C0267055">Erosive Esophagitis</annotation> in patients' xml_to_brat(sample_xml, 'patient_note.txt', 'patient_note.ann')
Step 3: Customization & Verification
- Use CUI as annotation type: If you want to label annotations with their CUI (e.g.,
C0267055instead ofannotation), modify the annotation line to:brat_annotations.append(f"T{annotation_id} {cui} {start_offset} {end_offset} {annotated_phrase}") - Check output: For your sample input, the script will generate:
patient_note.txt:Treatment of Erosive Esophagitis in patientspatient_note.ann:T1 annotation 14 33 Erosive Esophagitis
Which perfectly matches the BRAT format you provided.
Edge Cases to Keep In Mind
- Duplicate phrases: The script processes annotations in order, so even if the same text appears multiple times, each gets a unique ID and correct offsets.
- Multi-line text: Line breaks are counted as characters, so offsets will work even for wrapped or multi-paragraph XML input.
内容的提问来源于stack exchange,提问作者max
相关产品推荐
相关产品推荐

