You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python批量入库场景下,如何高效记录被跳过的样本?

优化跳过样本记录的几种方案

针对你要记录所有跳过样本(含整文件跳过的批量样本)的需求,以下是比“每个continue前调用自定义函数”更优的实现思路,兼顾可读性、可维护性和性能:


方案1:通用跳过记录器+细粒度异常捕获

核心是封装一个统一的记录函数,同时区分文件级跳过和样本级跳过,并在文件级跳过时尽可能关联其包含的所有样本(如果能安全获取的话)。同时替换裸except为具体异常捕获,避免隐藏未知问题。

示例代码

import logging
from datetime import datetime

def log_skipped(item_type, identifier, reason):
    """统一记录跳过的文件/样本"""
    # 基础日志输出
    log_msg = f"[{datetime.now()}] Skipped {item_type}: {identifier} | Reason: {reason}"
    logging.warning(log_msg)
    
    # 如果是文件级跳过,尝试获取并记录该文件下的所有样本
    if item_type == "file":
        try:
            samples = get_samples(identifier)
            for sample in samples:
                log_skipped("sample", sample, f"Parent file skipped: {reason}")
        except Exception as e:
            logging.warning(f"Failed to fetch samples from skipped file {identifier}: {str(e)}")

# 主流程
files = find_my_files()

for file in files:
    try:
        check_file(file)
    except Exception as e:
        log_skipped("file", file, f"Check failed: {str(e)}")
        continue
    
    try:
        process_file(file)
    except Exception as e:
        log_skipped("file", file, f"Processing failed: {str(e)}")
        continue
    
    try:
        samples = get_samples(file)
    except Exception as e:
        log_skipped("file", file, f"Failed to extract samples: {str(e)}")
        continue
    
    for sample in samples:
        try:
            check_sample(sample)
        except Exception as e:
            log_skipped("sample", sample, f"Check failed: {str(e)}")
            continue
        
        try:
            write_to_db(sample)
        except Exception as e:
            log_skipped("sample", sample, f"DB write failed: {str(e)}")
            continue

优势

  • 逻辑集中,修改记录规则只需调整log_skipped函数
  • 自动关联文件级跳过的所有样本,无需手动重复编写逻辑
  • 异常信息更详细,便于排查问题

方案2:批量收集跳过信息+统一落地

如果处理的样本量极大,实时写日志/数据库会有性能开销,可以先把所有跳过的信息存入内存列表,最后批量写入存储。

示例代码

import json
from datetime import datetime

# 存储所有跳过项的列表
skipped_records = []

def add_skipped(item_type, identifier, reason):
    """将跳过信息加入批量列表"""
    skipped_records.append({
        "type": item_type,
        "identifier": str(identifier),
        "reason": reason,
        "timestamp": datetime.now().isoformat()
    })

def save_skipped_records():
    """批量写入日志文件或数据库"""
    # 写入JSON文件示例
    with open("skipped_samples_log.json", "w", encoding="utf-8") as f:
        json.dump(skipped_records, f, indent=2, ensure_ascii=False)
    
    # 写入数据库示例(需自行实现)
    # for record in skipped_records:
    #     db.execute("INSERT INTO skipped_logs VALUES (?, ?, ?, ?)",
    #                (record["type"], record["identifier"], record["reason"], record["timestamp"]))

# 主流程(同方案1,仅把log_skipped替换为add_skipped)
files = find_my_files()

for file in files:
    try:
        check_file(file)
    except Exception as e:
        add_skipped("file", file, f"Check failed: {str(e)}")
        try:
            for sample in get_samples(file):
                add_skipped("sample", sample, f"Parent file skipped: Check failed")
        except Exception as e:
            add_skipped("file", file, f"Failed to fetch samples: {str(e)}")
        continue
    
    # ... 后续文件处理、样本处理逻辑,均使用add_skipped记录跳过项

# 所有文件处理完成后,批量落地记录
save_skipped_records()

优势

  • 减少IO操作次数,提升大样本量下的处理性能
  • 可以对跳过数据做后续分析(比如统计跳过原因占比)

方案3:代码重构+分层处理

通过函数拆分减少嵌套层级,把文件处理、样本处理的逻辑独立出来,在顶层统一处理跳过记录,让代码结构更清晰。

示例代码

def process_sample(sample):
    """处理单个样本,返回处理结果和失败原因"""
    try:
        check_sample(sample)
    except Exception as e:
        return False, f"Sample check failed: {str(e)}"
    
    try:
        write_to_db(sample)
    except Exception as e:
        return False, f"DB write failed: {str(e)}"
    
    return True, ""

def process_file(file):
    """处理单个文件,返回文件处理状态和跳过的样本列表"""
    try:
        check_file(file)
        process_file(file)
        samples = get_samples(file)
    except Exception as e:
        return False, f"File processing failed: {str(e)}"
    
    skipped_samples = []
    for sample in samples:
        success, reason = process_sample(sample)
        if not success:
            skipped_samples.append((sample, reason))
    
    return True, skipped_samples

# 主流程
files = find_my_files()
all_skipped = []

for file in files:
    file_ok, result = process_file(file)
    if not file_ok:
        # 文件级跳过,记录文件及所有关联样本
        all_skipped.append(("file", file, result))
        try:
            for sample in get_samples(file):
                all_skipped.append(("sample", sample, f"Parent file skipped: {result}"))
        except Exception as e:
            all_skipped.append(("file", file, f"Failed to fetch samples: {str(e)}"))
    else:
        # 样本级跳过,逐个记录
        for sample, reason in result:
            all_skipped.append(("sample", sample, reason))

# 输出所有跳过记录
for item_type, identifier, reason in all_skipped:
    logging.warning(f"Skipped {item_type}: {identifier} | {reason}")

优势

  • 代码模块化,每个函数只负责单一职责,便于调试和扩展
  • 跳过记录逻辑集中在顶层,避免分散在多个嵌套块中

内容的提问来源于stack exchange,提问作者Rash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 06:15:11