大TXT文件指定位置跳转及百余个眼动仪gaze数据文件计算需求
Hey there, let's tackle your two requirements one by one—both are totally manageable with Python scripting (it's perfect for efficient text processing and handling large files without loading everything into memory):
Large text files don't play nice with loading the entire content into memory, so we'll use targeted reading approaches:
按字节偏移量跳转:If you already know the exact byte position of your target (e.g., from a previous analysis), use
file.seek()to jump directly:def jump_to_byte_position(file_path, target_byte): with open(file_path, 'r') as f: f.seek(target_byte) # Read from this position onward content = f.read() print(content[:500]) # Print first 500 chars as a test按行号跳转:To jump to a specific line without reading the whole file, iterate through lines and count:
def jump_to_target_line(file_path, target_line_num): with open(file_path, 'r') as f: for current_line, line_content in enumerate(f, 1): if current_line == target_line_num: print(f"Reached line {target_line_num}:") print(line_content) # Continue reading subsequent lines if needed for line in f: print(line.strip()) break按内容标记跳转:Since your files split into calibration and test data, you can search for a unique marker to switch between sections:
def jump_to_test_data_section(file_path): with open(file_path, 'r') as f: for line in f: # Replace with your actual test data start marker if "Event: TestGaze" in line: print("Found test data starting here:") print(line.strip()) # Process all test data lines from this point for test_line in f: process_test_data_line(test_line) break
We'll split this into batch file traversal, calibration data processing, and test data processing:
Step 1: Traverse all gaze files
Use glob to easily grab all TXT files in your target directory:
import glob import os def process_all_gaze_files(root_dir): # Get all .txt files in the directory (and subdirs if needed) gaze_file_paths = glob.glob(os.path.join(root_dir, "*.txt")) for file_path in gaze_file_paths: print(f"Processing file: {os.path.basename(file_path)}") process_single_gaze_file(file_path)
Step 2: Process calibration data (limited variables)
Parse the calibration lines to extract key metrics and run calculations:
def process_calibration_line(line): # Parse line like: Event: Data - startTime 1563518990 endTime 1563619015 Gaze 885.638118989 316.57751978 line_parts = line.split() return { "start_time": int(line_parts[4]), "end_time": int(line_parts[6]), "gaze_x": float(line_parts[8]), "gaze_y": float(line_parts[9]) } def calculate_calibration_stats(calibration_entries): if not calibration_entries: return None # Example calculations: average gaze coordinates, total calibration duration avg_gaze_x = sum(entry["gaze_x"] for entry in calibration_entries) / len(calibration_entries) avg_gaze_y = sum(entry["gaze_y"] for entry in calibration_entries) / len(calibration_entries) total_duration = sum(entry["end_time"] - entry["start_time"] for entry in calibration_entries) return { "avg_gaze_x": round(avg_gaze_x, 2), "avg_gaze_y": round(avg_gaze_y, 2), "total_calib_duration": total_duration }
Step 3: Process test gaze data (more variables)
Adjust this parsing logic to match your actual test data format:
def process_test_data_line(line): # Replace with your test data's actual structure # Example line: Event: Test - Time 1563519020 Gaze 900.123 320.456 PupilLeft 4.5 PupilRight 4.3 FixationDuration 120 line_parts = line.split() return { "timestamp": int(line_parts[3]), "gaze_x": float(line_parts[5]), "gaze_y": float(line_parts[6]), "pupil_left": float(line_parts[8]), "pupil_right": float(line_parts[10]), "fixation_duration": int(line_parts[12]) } def calculate_test_data_stats(test_entries): if not test_entries: return None # Example calculations: average pupil size, total fixation time avg_pupil_size = (sum(entry["pupil_left"] for entry in test_entries) + sum(entry["pupil_right"] for entry in test_entries)) / (2 * len(test_entries)) total_fixation_time = sum(entry["fixation_duration"] for entry in test_entries) return { "avg_pupil_size": round(avg_pupil_size, 2), "total_fixation_time": total_fixation_time }
Step 4: Combine everything into a single file processor
def process_single_gaze_file(file_path): calibration_data = [] test_data = [] is_test_section = False with open(file_path, 'r') as f: for line in f: line = line.strip() if not line: continue # Check for calibration data marker if line.startswith("Event: Data -"): calib_entry = process_calibration_line(line) calibration_data.append(calib_entry) # Check for test data start marker elif line.startswith("Event: Test -"): is_test_section = True test_entry = process_test_data_line(line) test_data.append(test_entry) # Process remaining test data lines elif is_test_section: test_entry = process_test_data_line(line) test_data.append(test_entry) # Calculate and print stats calib_stats = calculate_calibration_stats(calibration_data) if calib_stats: print(f"Calibration Stats: {calib_stats}") test_stats = calculate_test_data_stats(test_data) if test_stats: print(f"Test Data Stats: {test_stats}") # Optional: Save processed data to CSV for further analysis # import pandas as pd # pd.DataFrame(calibration_data).to_csv(f"{file_path}_calib.csv", index=False) # pd.DataFrame(test_data).to_csv(f"{file_path}_test.csv", index=False)
To run everything
# Replace with your actual directory path process_all_gaze_files("/path/to/your/gaze/files")
内容的提问来源于stack exchange,提问作者Bart R

