You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3处理传感器CSV文件:文件名递增与多文件合并需求

Merge Sensor CSV Files by Sensor ID with Auto-Incrementing Filenames

Hey there! As someone new to Python 3, I’ll break this down into simple, actionable steps so you can merge those sensor CSV files without hassle. We’ll cover traversing your folders, grouping data by sensor, merging files, and handling duplicate filenames gracefully.

Step 1: Import Required Libraries

First, we’ll use os to navigate your folders and pandas to handle CSV data (if you don’t have pandas installed, run pip install pandas in your terminal):

import os
import pandas as pd

Step 2: Create a Helper Function for Auto-Incrementing Filenames

This function will check if a file exists, and if it does, append a number (like sensor_123.csv → sensor_123_1.csv) until it finds an unused name:

def get_unique_filename(base_path):
    if not os.path.exists(base_path):
        return base_path
    
    # Split the filename into name and extension
    name, ext = os.path.splitext(base_path)
    counter = 1
    
    # Keep incrementing until we find a unique filename
    while True:
        new_path = f"{name}_{counter}{ext}"
        if not os.path.exists(new_path):
            return new_path
        counter += 1

Step 3: Traverse Folders and Collect CSV Files

We’ll use os.walk() to go through every folder and subfolder, collecting paths to all .csv files:

def collect_csv_files(root_folder):
    csv_files = []
    for dirpath, _, filenames in os.walk(root_folder):
        for filename in filenames:
            if filename.lower().endswith(".csv"):
                csv_files.append(os.path.join(dirpath, filename))
    return csv_files

# Replace this with the path to your top-level folder containing all CSV folders
root_dir = "/path/to/your/main/folder"
all_csvs = collect_csv_files(root_dir)

Step 4: Group CSV Files by Sensor

How you group depends on how your sensor is identified. Let’s cover two common scenarios:

Scenario A: Sensor ID is in the Filename

If your CSV files are named like sensor_123_20240101.csv (where 123 is the sensor ID), we can extract the ID from the filename:

sensor_groups = {}

for csv_path in all_csvs:
    # Extract sensor ID from filename (adjust the split logic to match your naming)
    filename = os.path.basename(csv_path)
    # Example split: "sensor_123_20240101.csv" → split on "_" → ["sensor", "123", "20240101.csv"]
    sensor_id = filename.split("_")[1]
    
    if sensor_id not in sensor_groups:
        sensor_groups[sensor_id] = []
    sensor_groups[sensor_id].append(csv_path)

Scenario B: Sensor ID is a Column in the CSV

If each CSV has a column named sensor_id (or similar), we’ll read a small chunk of each file to get the sensor ID:

sensor_groups = {}

for csv_path in all_csvs:
    # Read just the first row to get the sensor ID (adjust column name if needed)
    df_sample = pd.read_csv(csv_path, nrows=1)
    sensor_id = str(df_sample["sensor_id"].iloc[0])
    
    if sensor_id not in sensor_groups:
        sensor_groups[sensor_id] = []
    sensor_groups[sensor_id].append(csv_path)

Step 5: Merge Each Sensor’s CSV Files

Now we’ll merge all CSV files for each sensor into a single file, using our unique filename function:

# Choose a folder to save merged files (create it if it doesn't exist)
output_dir = "/path/to/save/merged_files"
os.makedirs(output_dir, exist_ok=True)

for sensor_id, csv_paths in sensor_groups.items():
    # Create base filename for the merged sensor data
    base_filename = os.path.join(output_dir, f"sensor_{sensor_id}.csv")
    unique_filename = get_unique_filename(base_filename)
    
    # Initialize an empty list to store dataframes
    merged_data = []
    
    for idx, csv_path in enumerate(csv_paths):
        # Read the CSV (add encoding="utf-8" or "gbk" if you get encoding errors)
        df = pd.read_csv(csv_path)
        
        # Skip header for all files except the first one
        if idx > 0:
            df = df[1:]
        
        merged_data.append(df)
    
    # Combine all dataframes into one
    final_df = pd.concat(merged_data, ignore_index=True)
    
    # Save the merged file
    final_df.to_csv(unique_filename, index=False)
    print(f"Merged data for sensor {sensor_id} saved to {unique_filename}")

Notes for Large CSV Files

If your CSV files are extremely large (too big to fit in memory), modify the merging step to append to the output file incrementally instead of storing all data in memory:

for sensor_id, csv_paths in sensor_groups.items():
    base_filename = os.path.join(output_dir, f"sensor_{sensor_id}.csv")
    unique_filename = get_unique_filename(base_filename)
    
    for idx, csv_path in enumerate(csv_paths):
        with open(csv_path, "r", encoding="utf-8") as infile, open(unique_filename, "a", encoding="utf-8") as outfile:
            if idx > 0:
                # Skip header line for subsequent files
                next(infile)
            # Write all lines from infile to outfile
            for line in infile:
                outfile.write(line)
    print(f"Merged large files for sensor {sensor_id} saved to {unique_filename}")

This approach uses basic file I/O instead of pandas, which is more memory-efficient for huge files.


内容的提问来源于stack exchange,提问作者makerofmaps1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:36:47