You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量导入.cnv文件并提取指定文件名片段作为标识列

批量导入并合并CNV文件,添加文件名提取的标识列

你可以在读取每个CNV文件时,从文件名中提取指定片段作为新列,再合并所有DataFrame。以下是修改后的完整代码:

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import glob
import os
import re

# 获取目标文件夹下的所有指定CNV文件
path = '/Users/mariacristinaalvarez/Documents/High3/'
csv_files = glob.glob(path + "/*_high3.cnv")

# 定义函数:从文件名提取目标标识
def extract_id(filename):
    # 提取纯文件名(去除路径)
    base_name = os.path.basename(filename)
    # 用正则匹配'HLY2202_'之后、'_high3'之前的数字部分
    match = re.search(r'HLY2202_(\d+)_high3', base_name)
    if match:
        return match.group(1)
    return None  # 匹配失败时返回None

# 读取每个文件并添加标识列
df_list = []
for file in csv_files:
    # 按原参数读取文件
    df = pd.read_csv(file, encoding="ISO-8859-1", delim_whitespace=True, skiprows=316, header=None)
    # 新增列填充提取到的标识
    df['station_id'] = extract_id(file)
    df_list.append(df)

# 合并所有DataFrame并重置索引
stations_df = pd.concat(df_list, ignore_index=True)

关键细节说明:

  • 文件名提取逻辑:用正则表达式r'HLY2202_(\d+)_high3'精准定位目标片段,其中(\d+)会捕获连续数字(即示例中的"008"),适配你给出的文件名固定结构。
  • 修正原代码问题:原代码中csv_files1是笔误,改为定义好的csv_files;改用循环读取文件,更直观地给每个DataFrame添加新列。
  • 灵活适配:如果文件名前缀是HLYXXXX_这类可变数字格式,可将正则改为r'HLY\d+_(\d+)_high3',兼容更多相似命名的文件。

内容的提问来源于stack exchange,提问作者Maria Cristina Alvarez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 17:02:22