TDMS大文件加载进度条实现与加载性能优化咨询
Part 1: Adding a Progress Bar with tqdm or progressbar2
The key challenge with TdmsFile() is that the default full-file read doesn't expose progress updates. Instead, we need to read the file incrementally by first fetching its structure, then iterating through channels/groups while updating the progress bar.
Using tqdm (Recommended)
Here's a step-by-step implementation with the nptdms library (assuming that's what you're using):
from nptdms import TdmsFile from tqdm import tqdm import numpy as np # Open the TDMS file in read mode with TdmsFile.open("large_file.tdms") as tdms_file: # Get all groups and channels to calculate total work all_channels = [] for group in tdms_file.groups(): all_channels.extend(tdms_file[group.name].channels()) total_channels = len(all_channels) loaded_data = {} # Initialize tqdm progress bar (auto-calculates remaining time) with tqdm(total=total_channels, desc="Loading TDMS Channels") as pbar: for channel in all_channels: # Read channel data (adjust dtype if needed for memory efficiency) loaded_data[f"{channel.group.name}/{channel.name}"] = channel.read_data() # Update progress bar after each channel is loaded pbar.update(1) # Use loaded_data for your downstream tasks
If you prefer progressbar2, the logic is nearly identical—just swap out the tqdm block:
import progressbar with progressbar.ProgressBar(max_value=total_channels, widgets=[ progressbar.Percentage(), ' ', progressbar.Bar(), ' ', progressbar.ETA() ]) as pbar: for idx, channel in enumerate(all_channels): loaded_data[f"{channel.group.name}/{channel.name}"] = channel.read_data() pbar.update(idx + 1)
Why This Works
By breaking the load into individual channel reads, we can track each completed operation. Tqdm even estimates remaining time automatically based on the average speed of previous reads, which is perfect for long-running loads.
Part 2: Optimizing Loading Efficiency for Large TDMS Files
For files taking 5+ minutes to load (and repeated operations), here are actionable optimizations to speed things up without losing data:
1. Only Load What You Need
Avoid reading the entire file if you only need specific groups or channels. This cuts down on both IO and memory usage drastically:
with TdmsFile.open("large_file.tdms") as tdms_file: # Load only a specific group and its target channels target_group = tdms_file["MyCriticalGroup"] needed_channels = ["Temperature", "Pressure", "FlowRate"] loaded_data = { chan: target_group[chan].read_data() for chan in tqdm(needed_channels, desc="Loading Selected Channels") }
2. Use Memory Mapping
The nptdms library supports memory mapping with the memmap=True parameter. This maps the file directly to system memory instead of loading all data into RAM at once—ideal for extremely large files that would otherwise crash your session:
# Read a channel with memory mapping (acts like a numpy array but loads on-demand) channel_data = channel.read_data(memmap=True)
3. Parallelize Channel Reads
Since reading channels is mostly IO-bound, multi-threading can speed up the process by loading multiple channels at once. Here's how to do it with concurrent.futures:
from concurrent.futures import ThreadPoolExecutor with TdmsFile.open("large_file.tdms") as tdms_file: all_channels = [] for group in tdms_file.groups(): all_channels.extend(tdms_file[group.name].channels()) loaded_data = {} def load_channel(channel): key = f"{channel.group.name}/{channel.name}" loaded_data[key] = channel.read_data() # Use 4 threads (adjust based on your system's IO capacity) with ThreadPoolExecutor(max_workers=4) as executor: # Wrap executor.map with tqdm to track parallel progress list(tqdm(executor.map(load_channel, all_channels), total=len(all_channels), desc="Parallel Loading"))
4. Preprocess to a Faster Format
If you need to load the file multiple times, convert the necessary data to a more efficient format like Parquet or HDF5 once. These formats have faster read speeds, better compression, and are optimized for repeated access:
import pandas as pd # After loading needed channels into a dict, convert to DataFrame df = pd.DataFrame(loaded_data) # Save to Parquet (requires pyarrow or fastparquet) df.to_parquet("preprocessed_tdms_data.parquet") # Later, load in seconds instead of minutes: df = pd.read_parquet("preprocessed_tdms_data.parquet")
5. Upgrade nptdms
Make sure you're using the latest version of the library—developers regularly fix performance bottlenecks and add optimizations:
pip install --upgrade nptdms
内容的提问来源于stack exchange,提问作者Ivan Ezquerra

