You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从多文件中提取尺寸不固定的CC与FZ矩阵?

提取多文件中不同尺寸的CC/FZ标注矩阵方案

我来给你梳理一个通用的解决方案,不管矩阵是3x3、8x8还是1x1,都能自动识别并提取CC/FZ标注的矩阵,同时适配多文件场景。

核心思路拆解

  • 批量遍历文件:自动处理目标目录下的所有文件(也可以指定特定后缀的文件)
  • 精准定位标注:扫描文件内容,找到标记为"CC"或"FZ"的行,以此作为对应矩阵的起始信号
  • 自动识别矩阵边界:从标注行的下一行开始,连续读取有效数据行(直到遇到空行、新标注行或非数字内容),自动统计矩阵的行数和列数
  • 结构化存储结果:将每个文件中提取到的CC/FZ矩阵分别存储,方便后续查看或进一步处理

Python代码实现示例

import os

def extract_cc_fz_matrices(file_path):
    matrices = {}
    current_label = None
    current_matrix = []
    
    with open(file_path, 'r', encoding='utf-8') as f:
        for line in f:
            line = line.strip()
            # 跳过空行
            if not line:
                # 如果当前正在收集矩阵,说明矩阵结束
                if current_label and current_matrix:
                    matrices[current_label] = current_matrix
                    current_label = None
                    current_matrix = []
                continue
            
            # 识别CC/FZ标注行(可根据实际标注格式调整判断逻辑)
            if 'CC' in line:
                # 先保存上一个未完成的矩阵
                if current_label and current_matrix:
                    matrices[current_label] = current_matrix
                current_label = 'CC'
                current_matrix = []
                continue
            elif 'FZ' in line:
                if current_label and current_matrix:
                    matrices[current_label] = current_matrix
                current_label = 'FZ'
                current_matrix = []
                continue
            
            # 收集当前标注对应的矩阵行
            if current_label is not None:
                # 分割行内元素(假设用空格/制表符分隔)
                elements = line.split()
                # 转换为数值类型(不需要的话可以跳过此步)
                try:
                    elements = [float(num) for num in elements]
                except ValueError:
                    # 遇到非数字行,视为矩阵结束
                    matrices[current_label] = current_matrix
                    current_label = None
                    current_matrix = []
                    continue
                current_matrix.append(elements)
        
        # 处理文件末尾的最后一个矩阵
        if current_label and current_matrix:
            matrices[current_label] = current_matrix
    
    return matrices

def process_multiple_files(directory):
    all_results = {}
    # 遍历目录下所有文件(可添加后缀过滤,比如只处理.txt)
    for filename in os.listdir(directory):
        file_path = os.path.join(directory, filename)
        if os.path.isfile(file_path):
            print(f"正在处理文件: {filename}")
            file_matrices = extract_cc_fz_matrices(file_path)
            if file_matrices:
                all_results[filename] = file_matrices
    return all_results

# 使用示例
if __name__ == "__main__":
    target_dir = "./your_files_folder"  # 替换成你的文件目录路径
    results = process_multiple_files(target_dir)
    
    # 输出提取结果
    for filename, matrices in results.items():
        print(f"\n=== 文件 {filename} 提取结果 ===")
        for label, matrix in matrices.items():
            row_count = len(matrix)
            col_count = len(matrix[0]) if matrix else 0
            print(f"\n{label} 矩阵 ({row_count}x{col_count}):")
            for row in matrix:
                print(row)

关键细节调整说明

  • 标注匹配逻辑:如果你的标注格式是【CC】、Label: CC这类特殊格式,直接修改'CC' in line的判断条件即可,比如改成line.startswith("CC:")
  • 元素类型处理:如果矩阵是整数,把float(num)改成int(num);如果是字符串,直接去掉类型转换步骤
  • 文件过滤:想只处理特定后缀的文件,比如.txt,可以把os.listdir(directory)替换成glob.glob(f"{directory}/*.txt")
  • 异常处理:代码里加入了非数字行的判断,避免格式错误的内容干扰矩阵提取,你还可以添加文件读取失败的捕获逻辑

扩展优化建议

  • 加入日志记录,方便追踪哪个文件处理出了问题
  • 支持自定义标注关键词(把CC/FZ做成函数参数,灵活适配其他标注)
  • 增加矩阵格式校验,比如检查每行的元素数量是否一致,避免无效矩阵
  • 支持将提取结果导出为CSV、JSON等格式,方便后续分析使用

内容的提问来源于stack exchange,提问作者hsayya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:01:44