Python中如何创建二维列表数组以处理CSV数据计算列均值
Hey there! Let's break this down step by step—first we'll fix that one-dimensional list problem, then cover how to compare your two CSVs and calculate column means.
1. Convert to a 2D List (Rows as Sublists)
Your current code appends every single cell directly to my_list, which flattens everything into one big list. Instead, we need to create a sublist for each row, convert all cells in that row, then add the sublist to my_list.
Here's the revised code:
import csv def conv(s): try: # Convert numerical strings (including scientific notation) to float return float(s) except ValueError: # Return non-numerical values (like dates/times) as-is return s with open('Illia1.csv', newline='') as data: reader = csv.reader(data, delimiter=",") # Skip the first 8 header rows (cleaner than a while loop) for _ in range(8): next(reader, None) my_list = [] for row in reader: # Convert each cell in the row and store as a sublist converted_row = [conv(cell) for cell in row] my_list.append(converted_row) # Now my_list is 2D: each element is a row from the CSV print(my_list)
This will give you output like:
[ ['2001/01/24', '00:00.0', 1000.0, 1000.0, -0.0001, 0.0008], ['2001/01/25', '00:01.0', 2000.0, 2000.0, -0.0002, 0.0009], ... ]
2. Compare Two CSV Files for Differences
Since your two CSVs have different headers but identical data content, we'll process both into 2D lists first, then compare row by row (and cell by cell if needed).
Here's a function to handle the comparison:
def compare_csvs(file1_path, file2_path, skip_rows=8): # Helper to load CSV into 2D list def load_csv(file_path): with open(file_path, newline='') as f: reader = csv.reader(f, delimiter=",") for _ in range(skip_rows): next(reader, None) return [[conv(cell) for cell in row] for row in reader] data1 = load_csv(file1_path) data2 = load_csv(file2_path) differences = [] # Check row count first if len(data1) != len(data2): differences.append(f"Row count mismatch: {len(data1)} vs {len(data2)} rows") # Compare each row and cell for row_idx, (row1, row2) in enumerate(zip(data1, data2)): if len(row1) != len(row2): differences.append(f"Row {row_idx+1}: Column count mismatch ({len(row1)} vs {len(row2)})") continue for col_idx, (val1, val2) in enumerate(zip(row1, row2)): # For floats, use a small tolerance to avoid precision issues if isinstance(val1, float) and isinstance(val2, float): if abs(val1 - val2) > 1e-9: differences.append(f"Row {row_idx+1}, Column {col_idx+1}: {val1} vs {val2}") elif val1 != val2: differences.append(f"Row {row_idx+1}, Column {col_idx+1}: {val1} vs {val2}") return differences # Usage diff_results = compare_csvs('Illia1.csv', 'Illia2.csv') if diff_results: print("Found differences:") for diff in diff_results: print(f"- {diff}") else: print("No differences between the two CSVs.")
3. Calculate Column Means
Since the first two columns are dates/times, we'll skip those and calculate means for the numerical columns only:
def calculate_column_means(data): if not data: return [] means = [] # Start from column 2 (index 2) since first two are non-numerical for col_idx in range(2, len(data[0])): # Collect all float values in the column (skip any non-numeric entries) column_values = [row[col_idx] for row in data if isinstance(row[col_idx], float)] if column_values: mean = sum(column_values) / len(column_values) means.append(mean) else: means.append(None) # Mark columns with no numeric data return means # Usage column_means = calculate_column_means(my_list) print("Column means (starting from 3rd column):") for idx, mean in enumerate(column_means, start=3): print(f"Column {idx}: {mean}")
内容的提问来源于stack exchange,提问作者IluxaKuk

