You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python读取、拆分含4个Item的TextGrid文件并写入文件?

Split TextGrid Tiers into Separate Files Using Python

If you need to split a multi-tier TextGrid into individual files (each containing one tier), you have two reliable approaches: using a dedicated library (recommended for robustness) or writing a custom parser for simple, consistent cases.


praatio is a purpose-built Python library for working with TextGrid and speech annotation data. It handles all parsing nuances automatically, so you don’t have to worry about edge cases like point tiers or varying formatting.

Step 1: Install praatio

First, install the library via pip:

pip install praatio

Step 2: Code to Split Tiers

This script loads your TextGrid, extracts each tier, and saves each as a standalone TextGrid file:

from praatio import textgrid

# Load the original TextGrid (preserve empty intervals)
tg = textgrid.openTextgrid("your_input_file.TextGrid", includeEmptyIntervals=True)

# Iterate over each tier and save it to a separate file
for tier_name in tg.tierNames:
    # Create a new TextGrid with the same time range as the original
    new_tg = textgrid.Textgrid(xmin=tg.xmin, xmax=tg.xmax)
    # Add the current tier to the new TextGrid
    new_tg.addTier(tg.getTier(tier_name))
    
    # Generate a descriptive filename (matches your item [1]-[4] numbering)
    tier_index = tg.tierNames.index(tier_name) + 1
    output_filename = f"tier_{tier_index}_{tier_name}.TextGrid"
    
    # Save the new TextGrid
    new_tg.save(output_filename, format="short_textgrid", includeBlankSpaces=True)

Explanation:

  • openTextgrid: Loads your original file while preserving empty intervals (critical for maintaining timing accuracy).
  • Textgrid(xmin=..., xmax=...): Ensures the new file uses the same time range as your original TextGrid.
  • addTier: Adds the single extracted tier to the new empty TextGrid.
  • save: Writes the file with a clear, descriptive name that includes the tier’s original index and name.

Approach 2: Custom Parser (For Simple Cases)

If you can’t install external libraries, this basic parser works for standard IntervalTiers with consistent formatting (like your example):

Step 1: Code for Custom Parsing

def split_textgrid(input_path):
    # Read all lines from the input file
    with open(input_path, 'r', encoding='utf-8') as f:
        lines = [line.rstrip('\n') for line in f]
    
    # Locate the end of the header section
    header_end_idx = None
    for i, line in enumerate(lines):
        if line.strip() == "item []:":
            header_end_idx = i
            break
    
    if header_end_idx is None:
        raise ValueError("Invalid TextGrid format: Could not find item []: section")
    
    # Modify the header to reflect a single tier (change size from 4 to 1)
    modified_header = []
    for line in lines[:header_end_idx]:
        if line.strip().startswith("size ="):
            modified_header.append("size = 1")
        else:
            modified_header.append(line)
    modified_header.append("item []:")  # Re-add the item section marker
    
    # Split the file into individual tier blocks
    items = []
    current_item = []
    in_item = False
    
    for line in lines[header_end_idx+1:]:
        if line.strip().startswith("item ["):
            if in_item:
                items.append(current_item)
            current_item = [line]
            in_item = True
        else:
            if in_item:
                current_item.append(line)
    
    # Add the final tier block
    if in_item:
        items.append(current_item)
    
    # Write each tier to a separate file
    for idx, item in enumerate(items, 1):
        # Extract the tier name for the filename
        tier_name = None
        for line in item:
            if line.strip().startswith("name ="):
                tier_name = line.split('=')[1].strip().strip('"')
                break
        if tier_name is None:
            tier_name = f"unnamed_tier_{idx}"
        
        output_filename = f"tier_{idx}_{tier_name}.TextGrid"
        with open(output_filename, 'w', encoding='utf-8') as f:
            f.write('\n'.join(modified_header) + '\n')
            f.write('\n'.join(item) + '\n')

# Run the function with your input file
split_textgrid("your_input_file.TextGrid")

Explanation:

  • Header Modification: We adjust the size parameter from 4 to 1 to indicate each output file has one tier, while keeping all other header details (like time range) intact.
  • Tier Splitting: We iterate through the file to separate each item [n]: block into its own list.
  • File Writing: Each tier is combined with the modified header and saved to a file with a clear name.

Key Notes:

  • The praatio approach is strongly recommended because it handles all TextGrid variations (point tiers, non-standard interval formatting, etc.) that a custom parser might miss.
  • Replace "your_input_file.TextGrid" with the actual path to your TextGrid file before running either script.

内容的提问来源于stack exchange,提问作者Qsec

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:54:01