You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多CSV文件读取与平均响应时间计算问题求助

嘿,作为Python新手,处理多文件确实容易卡壳,我来一步步帮你解决这个问题!

核心思路

你的需求本质是批量读取文件→提取目标列→按条件聚合数据→计算均值,我们可以分层次来实现:先按参与者分组处理文件,再在每个参与者的文件中收集对应条件的响应时间,最后计算平均值。

基础版:用内置csv模块实现(适合理解原理)

这个版本不用额外安装库,适合新手理解每一步的逻辑:

from pathlib import Path
import csv

# 第一步:设置你的数据根目录(改成你实际的文件夹路径)
data_root = Path("participants_data")

# 第二步:遍历每个参与者的文件夹
for participant_folder in data_root.iterdir():
    if not participant_folder.is_dir():
        continue  # 跳过不是文件夹的文件
    participant_name = participant_folder.name
    print(f"正在处理参与者: {participant_name}")
    
    # 初始化字典,用来存储每个条件的所有响应时间
    condition_times = {
        "condition_A": [],
        "condition_B": [],
        "condition_C": []
        # 注意:这里要改成你文件中实际的条件名称
    }
    
    # 第三步:遍历当前参与者的所有CSV文件
    for csv_file in participant_folder.glob("*.csv"):
        with open(csv_file, "r", newline="", encoding="utf-8") as f:
            reader = csv.reader(f)
            
            # 如果你的CSV文件有表头,先跳过表头行(解开下面这行注释)
            # next(reader)
            
            for row in reader:
                # 提取第3列(条件,索引2,因为Python从0开始计数)和第6列(响应时间,索引5)
                condition = row[2]
                try:
                    response_time = float(row[5])  # 把文本转成数值型
                    # 将响应时间添加到对应条件的列表中
                    if condition in condition_times:
                        condition_times[condition].append(response_time)
                    else:
                        print(f"⚠️ 警告:文件{csv_file}中发现未知条件{condition},已跳过")
                except ValueError:
                    print(f"⚠️ 警告:文件{csv_file}中某行的响应时间不是有效数字,已跳过")
    
    # 第四步:计算并输出每个条件的平均响应时间
    print(f"{participant_name} 的平均响应时间:")
    for condition, times in condition_times.items():
        if times:  # 确保该条件有有效数据
            avg_rt = sum(times) / len(times)
            print(f"  {condition}: {avg_rt:.2f} ms")
        else:
            print(f"  {condition}: 无有效响应时间数据")
    print("------------------------")
进阶版:用pandas简化操作(高效推荐)

如果你的数据量较大,或者想更高效地处理,pandas会帮你省很多代码:

首先需要安装pandas(如果没装的话):

pip install pandas

然后运行以下代码:

import pandas as pd
from pathlib import Path

data_root = Path("participants_data")

for participant_folder in data_root.iterdir():
    if not participant_folder.is_dir():
        continue
    participant_name = participant_folder.name
    print(f"正在处理参与者: {participant_name}")
    
    # 收集当前参与者的所有CSV文件路径
    all_csv_files = list(participant_folder.glob("*.csv"))
    if not all_csv_files:
        print("  ❌ 未找到任何CSV文件")
        continue
    
    # 批量读取文件,只提取第3列和第6列(索引2和5)
    df_list = []
    for file in all_csv_files:
        # names参数给列命名,header=None表示文件没有表头(如果有表头,改成header=0)
        df = pd.read_csv(file, usecols=[2, 5], names=["condition", "response_time"], header=None)
        df_list.append(df)
    
    # 合并所有文件的数据到一个DataFrame
    combined_data = pd.concat(df_list, ignore_index=True)
    
    # 按条件分组计算平均响应时间
    avg_response_times = combined_data.groupby("condition")["response_time"].mean()
    
    # 输出结果(保留两位小数)
    print(f"{participant_name} 的平均响应时间:")
    print(avg_response_times.round(2))
    print("------------------------")
实用小贴士
  1. 确认分隔符:如果你的CSV文件是用制表符(\t)或其他符号分隔的,要在读取时指定:
    • 基础版:csv.reader(f, delimiter="\t")
    • pandas版:pd.read_csv(file, sep="\t", ...)
  2. 表头处理:如果文件第一行是表头(比如"条件"、"响应时间"),记得调整代码跳过表头(基础版用next(reader),pandas版用header=0)
  3. 保存结果:可以把每个参与者的平均结果保存成新的CSV文件,比如用pandas的avg_response_times.to_csv(f"{participant_name}_avg_rt.csv")

内容的提问来源于stack exchange,提问作者user9524759

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:13:28