You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flask上传文件两次计算MD5哈希值结果不一致问题求助

Flask文件上传MD5哈希计算不一致的解决方案

问题现象

在Flask文件上传接口中,通过两次MD5哈希校验判断文件是否存在时,出现以下异常:

  • 同一文件两次计算的哈希值完全不同
  • 连续调用两次哈希计算,结果也不一致
  • 后续上传其他文件时,哈希值会复用最后一次的错误结果

相关代码如下:

路由代码

@app.route("/file_upload", methods=['POST'])
def file_upload():
uploaded_files = request.files.getlist("pdf")
for file in uploaded_files:
    logger.info(f"Processing item {file.filename}")
    # Check if file is already in database
    file_hash = md5_it(file.read())  # !! FIRST HASH
    print(file_hash)
    in_db = get_file(file_hash)
    if in_db:
        logger.info(f"File already in the database, skipping...")
        # File already in database, skipping OCR
        print(in_db)
        text_pdf = in_db.text
        pass
    else:
        logger.debug("Doing OCR...")
        ocred_pdf = ocr_pdf(file)

        logger.debug("Extracting text...")
        text_pdf = read_pdf(file.filename, ocred_pdf)

        logger.debug("Saving to DB")
        create_file(file.read(), ocred_pdf, file.filename, text_pdf) # !! SECOND HASH

return text_pdf

哈希函数代码

def md5_it(file):
    return hashlib.md5(file).hexdigest()

原因分析

核心问题是文件对象的指针位置偏移:

  • 调用file.read()后,文件指针会直接移动到文件末尾
  • 后续再次调用file.read()时,会返回空字节串b''
  • 基于空内容计算的MD5哈希是固定值(d41d8cd98f00b204e9800998ecf8427e),导致后续所有哈希计算都复用这个错误结果

解决步骤

有两种可靠的修复方式:

方式1:一次性读取文件内容到内存

先把文件完整内容读取到内存变量中,后续所有操作都基于这个变量,避免多次操作文件指针:

@app.route("/file_upload", methods=['POST'])
def file_upload():
    uploaded_files = request.files.getlist("pdf")
    for file in uploaded_files:
        logger.info(f"Processing item {file.filename}")
        # 一次性读取文件内容到内存
        file_content = file.read()
        # 使用内存中的内容计算哈希
        file_hash = md5_it(file_content)  
        print(file_hash)
        in_db = get_file(file_hash)
        if in_db:
            logger.info(f"File already in the database, skipping...")
            print(in_db)
            text_pdf = in_db.text
        else:
            logger.debug("Doing OCR...")
            # 如果ocr_pdf需要读取文件内容,先重置指针
            file.seek(0)
            ocred_pdf = ocr_pdf(file)

            logger.debug("Extracting text...")
            text_pdf = read_pdf(file.filename, ocred_pdf)

            logger.debug("Saving to DB")
            # 直接传入内存中的文件内容,无需再次read()
            create_file(file_content, ocred_pdf, file.filename, text_pdf) 
    return text_pdf

方式2:每次读取后重置文件指针

如果必须多次读取文件对象,每次读取后调用file.seek(0)将指针重置到文件开头:

@app.route("/file_upload", methods=['POST'])
def file_upload():
    uploaded_files = request.files.getlist("pdf")
    for file in uploaded_files:
        logger.info(f"Processing item {file.filename}")
        # 第一次读取并计算哈希
        file_hash = md5_it(file.read())  
        print(file_hash)
        # 重置指针到开头
        file.seek(0)
        in_db = get_file(file_hash)
        if in_db:
            logger.info(f"File already in the database, skipping...")
            print(in_db)
            text_pdf = in_db.text
        else:
            logger.debug("Doing OCR...")
            ocred_pdf = ocr_pdf(file)

            logger.debug("Extracting text...")
            text_pdf = read_pdf(file.filename, ocred_pdf)

            logger.debug("Saving to DB")
            # 再次重置指针后读取文件内容
            file.seek(0)
            create_file(file.read(), ocred_pdf, file.filename, text_pdf) 
    return text_pdf

注意事项

  • 如果ocr_pdf函数内部会调用file.read(),必须确保调用前指针处于文件开头位置
  • 优先选择方式1,减少重复IO操作,提升接口性能

内容的提问来源于stack exchange,提问作者Mike

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 10:33:17