You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取大型TDMS数据集避免内存溢出的正确方法

Fixing TDMS Memory Bloat When Reading & Plotting in Python 2.7

Hey there, let's break down why you're hitting memory issues and fix this step by step.

First, let's do the math to understand the problem: each of your 3.3e7-point arrays (if stored as 64-bit floats) takes up ~264MB of memory. Multiply that by 8 channels, and you're looking at over 2GB of raw data—before even accounting for Python's memory overhead, Ubuntu system usage, and IPython Notebook's background processes. With only 8GB of RAM, it's no surprise things get tight.

Here are the practical fixes to avoid loading the entire file into memory:

1. Use a TDMS Library That Supports On-Demand Reading

The nptdms library (compatible with Python 2.7) lets you read only the channels you need, instead of loading the entire file. First, install a Python 2.7-compatible version:

pip install nptdms==0.15.0

Example: Read Only the Channel You Want to Plot

Instead of loading all 8 channels, target just the one you need to visualize:

from nptdms import TdmsFile

# Open the file without preloading all data
with TdmsFile.open("your_dataset.tdms") as tdms_file:
    # Replace "Your Group" and "Your Channel" with your actual group/channel names
    target_channel = tdms_file["Your Group"]["Your Channel"]
    # Load only this channel's data into memory
    channel_data = target_channel[:]

This cuts your memory usage from 2GB+ to ~264MB—way more manageable for your 8GB RAM.

2. Downsample Data for Faster Plotting

3.3e7 points is way more than matplotlib needs to render a readable plot. Downsample by picking every Nth point to reduce the dataset size:

import matplotlib.pyplot as plt

# Downsample to 1/10th the original size (adjust the step as needed)
downsampled_data = channel_data[::10]

# Plot the downsampled data
plt.plot(downsampled_data, linewidth=0.5)
plt.title("Downsampled TDMS Channel Data")
plt.show()

This reduces the data to 3.3e6 points, which is trivial for matplotlib and uses only ~26MB of memory for the plot data.

3. Read & Plot in Chunks (For Extra Large Single Channels)

If even a single channel's full data is too big, read and plot in chunks to avoid loading everything at once:

fig, ax = plt.subplots()
chunk_size = 1_000_000  # Process 1 million points at a time
total_points = len(target_channel)

for start_idx in range(0, total_points, chunk_size):
    end_idx = min(start_idx + chunk_size, total_points)
    # Read just this chunk of data
    chunk = target_channel[start_idx:end_idx]
    # Plot the chunk with correct x-axis positioning
    ax.plot(range(start_idx, end_idx), chunk, linewidth=0.3)

plt.title("Chunked TDMS Channel Plot")
plt.show()

4. Clean Up Unused Memory Manually

IPython Notebook can hang onto unused data in memory. After loading/plotting, explicitly free up space:

import gc

# Delete variables you no longer need
del channel_data, downsampled_data
# Force garbage collection to reclaim memory
gc.collect()

内容的提问来源于stack exchange,提问作者nandhos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:30:20