如何用Python读取、拆分含4个Item的TextGrid文件并写入文件?
If you need to split a multi-tier TextGrid into individual files (each containing one tier), you have two reliable approaches: using a dedicated library (recommended for robustness) or writing a custom parser for simple, consistent cases.
Approach 1: Use the praatio Library (Recommended)
praatio is a purpose-built Python library for working with TextGrid and speech annotation data. It handles all parsing nuances automatically, so you don’t have to worry about edge cases like point tiers or varying formatting.
Step 1: Install praatio
First, install the library via pip:
pip install praatio
Step 2: Code to Split Tiers
This script loads your TextGrid, extracts each tier, and saves each as a standalone TextGrid file:
from praatio import textgrid # Load the original TextGrid (preserve empty intervals) tg = textgrid.openTextgrid("your_input_file.TextGrid", includeEmptyIntervals=True) # Iterate over each tier and save it to a separate file for tier_name in tg.tierNames: # Create a new TextGrid with the same time range as the original new_tg = textgrid.Textgrid(xmin=tg.xmin, xmax=tg.xmax) # Add the current tier to the new TextGrid new_tg.addTier(tg.getTier(tier_name)) # Generate a descriptive filename (matches your item [1]-[4] numbering) tier_index = tg.tierNames.index(tier_name) + 1 output_filename = f"tier_{tier_index}_{tier_name}.TextGrid" # Save the new TextGrid new_tg.save(output_filename, format="short_textgrid", includeBlankSpaces=True)
Explanation:
openTextgrid: Loads your original file while preserving empty intervals (critical for maintaining timing accuracy).Textgrid(xmin=..., xmax=...): Ensures the new file uses the same time range as your original TextGrid.addTier: Adds the single extracted tier to the new empty TextGrid.save: Writes the file with a clear, descriptive name that includes the tier’s original index and name.
Approach 2: Custom Parser (For Simple Cases)
If you can’t install external libraries, this basic parser works for standard IntervalTiers with consistent formatting (like your example):
Step 1: Code for Custom Parsing
def split_textgrid(input_path): # Read all lines from the input file with open(input_path, 'r', encoding='utf-8') as f: lines = [line.rstrip('\n') for line in f] # Locate the end of the header section header_end_idx = None for i, line in enumerate(lines): if line.strip() == "item []:": header_end_idx = i break if header_end_idx is None: raise ValueError("Invalid TextGrid format: Could not find item []: section") # Modify the header to reflect a single tier (change size from 4 to 1) modified_header = [] for line in lines[:header_end_idx]: if line.strip().startswith("size ="): modified_header.append("size = 1") else: modified_header.append(line) modified_header.append("item []:") # Re-add the item section marker # Split the file into individual tier blocks items = [] current_item = [] in_item = False for line in lines[header_end_idx+1:]: if line.strip().startswith("item ["): if in_item: items.append(current_item) current_item = [line] in_item = True else: if in_item: current_item.append(line) # Add the final tier block if in_item: items.append(current_item) # Write each tier to a separate file for idx, item in enumerate(items, 1): # Extract the tier name for the filename tier_name = None for line in item: if line.strip().startswith("name ="): tier_name = line.split('=')[1].strip().strip('"') break if tier_name is None: tier_name = f"unnamed_tier_{idx}" output_filename = f"tier_{idx}_{tier_name}.TextGrid" with open(output_filename, 'w', encoding='utf-8') as f: f.write('\n'.join(modified_header) + '\n') f.write('\n'.join(item) + '\n') # Run the function with your input file split_textgrid("your_input_file.TextGrid")
Explanation:
- Header Modification: We adjust the
sizeparameter from 4 to 1 to indicate each output file has one tier, while keeping all other header details (like time range) intact. - Tier Splitting: We iterate through the file to separate each
item [n]:block into its own list. - File Writing: Each tier is combined with the modified header and saved to a file with a clear name.
Key Notes:
- The
praatioapproach is strongly recommended because it handles all TextGrid variations (point tiers, non-standard interval formatting, etc.) that a custom parser might miss. - Replace
"your_input_file.TextGrid"with the actual path to your TextGrid file before running either script.
内容的提问来源于stack exchange,提问作者Qsec

