Python 3处理传感器CSV文件:文件名递增与多文件合并需求
Hey there! As someone new to Python 3, I’ll break this down into simple, actionable steps so you can merge those sensor CSV files without hassle. We’ll cover traversing your folders, grouping data by sensor, merging files, and handling duplicate filenames gracefully.
Step 1: Import Required Libraries
First, we’ll use os to navigate your folders and pandas to handle CSV data (if you don’t have pandas installed, run pip install pandas in your terminal):
import os import pandas as pd
Step 2: Create a Helper Function for Auto-Incrementing Filenames
This function will check if a file exists, and if it does, append a number (like sensor_123.csv → sensor_123_1.csv) until it finds an unused name:
def get_unique_filename(base_path): if not os.path.exists(base_path): return base_path # Split the filename into name and extension name, ext = os.path.splitext(base_path) counter = 1 # Keep incrementing until we find a unique filename while True: new_path = f"{name}_{counter}{ext}" if not os.path.exists(new_path): return new_path counter += 1
Step 3: Traverse Folders and Collect CSV Files
We’ll use os.walk() to go through every folder and subfolder, collecting paths to all .csv files:
def collect_csv_files(root_folder): csv_files = [] for dirpath, _, filenames in os.walk(root_folder): for filename in filenames: if filename.lower().endswith(".csv"): csv_files.append(os.path.join(dirpath, filename)) return csv_files # Replace this with the path to your top-level folder containing all CSV folders root_dir = "/path/to/your/main/folder" all_csvs = collect_csv_files(root_dir)
Step 4: Group CSV Files by Sensor
How you group depends on how your sensor is identified. Let’s cover two common scenarios:
Scenario A: Sensor ID is in the Filename
If your CSV files are named like sensor_123_20240101.csv (where 123 is the sensor ID), we can extract the ID from the filename:
sensor_groups = {} for csv_path in all_csvs: # Extract sensor ID from filename (adjust the split logic to match your naming) filename = os.path.basename(csv_path) # Example split: "sensor_123_20240101.csv" → split on "_" → ["sensor", "123", "20240101.csv"] sensor_id = filename.split("_")[1] if sensor_id not in sensor_groups: sensor_groups[sensor_id] = [] sensor_groups[sensor_id].append(csv_path)
Scenario B: Sensor ID is a Column in the CSV
If each CSV has a column named sensor_id (or similar), we’ll read a small chunk of each file to get the sensor ID:
sensor_groups = {} for csv_path in all_csvs: # Read just the first row to get the sensor ID (adjust column name if needed) df_sample = pd.read_csv(csv_path, nrows=1) sensor_id = str(df_sample["sensor_id"].iloc[0]) if sensor_id not in sensor_groups: sensor_groups[sensor_id] = [] sensor_groups[sensor_id].append(csv_path)
Step 5: Merge Each Sensor’s CSV Files
Now we’ll merge all CSV files for each sensor into a single file, using our unique filename function:
# Choose a folder to save merged files (create it if it doesn't exist) output_dir = "/path/to/save/merged_files" os.makedirs(output_dir, exist_ok=True) for sensor_id, csv_paths in sensor_groups.items(): # Create base filename for the merged sensor data base_filename = os.path.join(output_dir, f"sensor_{sensor_id}.csv") unique_filename = get_unique_filename(base_filename) # Initialize an empty list to store dataframes merged_data = [] for idx, csv_path in enumerate(csv_paths): # Read the CSV (add encoding="utf-8" or "gbk" if you get encoding errors) df = pd.read_csv(csv_path) # Skip header for all files except the first one if idx > 0: df = df[1:] merged_data.append(df) # Combine all dataframes into one final_df = pd.concat(merged_data, ignore_index=True) # Save the merged file final_df.to_csv(unique_filename, index=False) print(f"Merged data for sensor {sensor_id} saved to {unique_filename}")
Notes for Large CSV Files
If your CSV files are extremely large (too big to fit in memory), modify the merging step to append to the output file incrementally instead of storing all data in memory:
for sensor_id, csv_paths in sensor_groups.items(): base_filename = os.path.join(output_dir, f"sensor_{sensor_id}.csv") unique_filename = get_unique_filename(base_filename) for idx, csv_path in enumerate(csv_paths): with open(csv_path, "r", encoding="utf-8") as infile, open(unique_filename, "a", encoding="utf-8") as outfile: if idx > 0: # Skip header line for subsequent files next(infile) # Write all lines from infile to outfile for line in infile: outfile.write(line) print(f"Merged large files for sensor {sensor_id} saved to {unique_filename}")
This approach uses basic file I/O instead of pandas, which is more memory-efficient for huge files.
内容的提问来源于stack exchange,提问作者makerofmaps1

