You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多进程处理多被试三维数据集的代码优化需求

Hey there! Let's tweak your code to process multiple subjects one after another—each subject will complete its full channel feature extraction via multiprocessing before we move on to the next. I'll also fix a redundant file-loading issue in your original code to make things more efficient!

Solution: Sequential Multiprocessing for Multiple Subjects

Here's the revised code that handles all your subjects automatically, with optimized data handling:

import numpy as np
import time
from multiprocessing import Pool

def cal_feature(args):
    data, ch_start = args
    # Calculate mean across the last dimension for the 8-channel chunk
    return np.mean(data[:, ch_start:ch_start+8, :], axis=-1)

if __name__ == '__main__':
    # List of subject IDs to process (adjust this to match your data files)
    sub_list = [1, 2, 3, 4]
    total_start = time.time()

    for sub in sub_list:
        print(f"Starting processing for subject {sub}...")
        sub_start = time.time()
        
        # Load the subject's data ONCE in the main process (no redundant reads!)
        data = np.load(f'data_{sub}.npy')
        
        # Generate channel start positions: 0, 8, 16, ..., 56 (8 chunks total)
        ch_starts = list(range(0, 64, 8))
        
        # Use multiprocessing to process each channel chunk in parallel
        with Pool(8) as p:
            # Pass pre-loaded data + channel start to each worker
            result = p.map(cal_feature, [(data, ch) for ch in ch_starts])
        
        # Combine chunk results and save the final feature matrix
        combined_result = np.concatenate(result, axis=1)
        np.save(f'subject_0_{sub}', combined_result)
        
        print(f"Subject {sub} finished! Time taken: {time.time() - sub_start:.2f} seconds")
    
    print(f"All subjects processed! Total elapsed time: {time.time() - total_start:.2f} seconds")

Key Improvements & Explanations

  • Sequential Subject Handling: We loop through each subject in sub_list—each subject goes through full processing (load → extract features → save) before the next one starts, just as you wanted.
  • Optimized Data Loading: The main process loads each subject's data once, instead of having every worker process reload the same file (your original code loaded the same file 8 times per subject!). This cuts down on unnecessary disk I/O and speeds things up.
  • Adjusted Feature Function: cal_feature now takes a tuple of (pre-loaded data, channel start index) so workers can directly operate on the existing data array.
  • Clean Result Saving: We concatenate the 8 chunk results along axis 1 to get a full (100, 64) feature matrix for each subject, then save it with the correct filename.

Bonus: Auto-Detect Subject Files

If you don't want to manually list subject IDs, you can auto-detect all data_*.npy files using glob:

import glob
# Get all matching data files and extract subject IDs
sub_files = glob.glob('data_*.npy')
sub_list = [int(file.split('_')[1].split('.')[0]) for file in sub_files]
# Sort to process subjects in numerical order
sub_list.sort()

内容的提问来源于stack exchange,提问作者Talha Anwar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:21:42