You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何创建二维列表数组以处理CSV数据计算列均值

Fixing the 1D List Issue & Next Steps for CSV Comparison + Mean Calculation

Hey there! Let's break this down step by step—first we'll fix that one-dimensional list problem, then cover how to compare your two CSVs and calculate column means.

1. Convert to a 2D List (Rows as Sublists)

Your current code appends every single cell directly to my_list, which flattens everything into one big list. Instead, we need to create a sublist for each row, convert all cells in that row, then add the sublist to my_list.

Here's the revised code:

import csv

def conv(s):
    try:
        # Convert numerical strings (including scientific notation) to float
        return float(s)
    except ValueError:
        # Return non-numerical values (like dates/times) as-is
        return s

with open('Illia1.csv', newline='') as data:
    reader = csv.reader(data, delimiter=",")
    # Skip the first 8 header rows (cleaner than a while loop)
    for _ in range(8):
        next(reader, None)
    
    my_list = []
    for row in reader:
        # Convert each cell in the row and store as a sublist
        converted_row = [conv(cell) for cell in row]
        my_list.append(converted_row)

# Now my_list is 2D: each element is a row from the CSV
print(my_list)

This will give you output like:

[
    ['2001/01/24', '00:00.0', 1000.0, 1000.0, -0.0001, 0.0008],
    ['2001/01/25', '00:01.0', 2000.0, 2000.0, -0.0002, 0.0009],
    ...
]

2. Compare Two CSV Files for Differences

Since your two CSVs have different headers but identical data content, we'll process both into 2D lists first, then compare row by row (and cell by cell if needed).

Here's a function to handle the comparison:

def compare_csvs(file1_path, file2_path, skip_rows=8):
    # Helper to load CSV into 2D list
    def load_csv(file_path):
        with open(file_path, newline='') as f:
            reader = csv.reader(f, delimiter=",")
            for _ in range(skip_rows):
                next(reader, None)
            return [[conv(cell) for cell in row] for row in reader]
    
    data1 = load_csv(file1_path)
    data2 = load_csv(file2_path)
    
    differences = []
    # Check row count first
    if len(data1) != len(data2):
        differences.append(f"Row count mismatch: {len(data1)} vs {len(data2)} rows")
    
    # Compare each row and cell
    for row_idx, (row1, row2) in enumerate(zip(data1, data2)):
        if len(row1) != len(row2):
            differences.append(f"Row {row_idx+1}: Column count mismatch ({len(row1)} vs {len(row2)})")
            continue
        for col_idx, (val1, val2) in enumerate(zip(row1, row2)):
            # For floats, use a small tolerance to avoid precision issues
            if isinstance(val1, float) and isinstance(val2, float):
                if abs(val1 - val2) > 1e-9:
                    differences.append(f"Row {row_idx+1}, Column {col_idx+1}: {val1} vs {val2}")
            elif val1 != val2:
                differences.append(f"Row {row_idx+1}, Column {col_idx+1}: {val1} vs {val2}")
    
    return differences

# Usage
diff_results = compare_csvs('Illia1.csv', 'Illia2.csv')
if diff_results:
    print("Found differences:")
    for diff in diff_results:
        print(f"- {diff}")
else:
    print("No differences between the two CSVs.")

3. Calculate Column Means

Since the first two columns are dates/times, we'll skip those and calculate means for the numerical columns only:

def calculate_column_means(data):
    if not data:
        return []
    
    means = []
    # Start from column 2 (index 2) since first two are non-numerical
    for col_idx in range(2, len(data[0])):
        # Collect all float values in the column (skip any non-numeric entries)
        column_values = [row[col_idx] for row in data if isinstance(row[col_idx], float)]
        if column_values:
            mean = sum(column_values) / len(column_values)
            means.append(mean)
        else:
            means.append(None)  # Mark columns with no numeric data
    
    return means

# Usage
column_means = calculate_column_means(my_list)
print("Column means (starting from 3rd column):")
for idx, mean in enumerate(column_means, start=3):
    print(f"Column {idx}: {mean}")

内容的提问来源于stack exchange,提问作者IluxaKuk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:26:33