Flask上传文件两次计算MD5哈希值结果不一致问题求助
Flask文件上传MD5哈希计算不一致的解决方案
问题现象
在Flask文件上传接口中,通过两次MD5哈希校验判断文件是否存在时,出现以下异常:
- 同一文件两次计算的哈希值完全不同
- 连续调用两次哈希计算,结果也不一致
- 后续上传其他文件时,哈希值会复用最后一次的错误结果
相关代码如下:
路由代码
@app.route("/file_upload", methods=['POST']) def file_upload(): uploaded_files = request.files.getlist("pdf") for file in uploaded_files: logger.info(f"Processing item {file.filename}") # Check if file is already in database file_hash = md5_it(file.read()) # !! FIRST HASH print(file_hash) in_db = get_file(file_hash) if in_db: logger.info(f"File already in the database, skipping...") # File already in database, skipping OCR print(in_db) text_pdf = in_db.text pass else: logger.debug("Doing OCR...") ocred_pdf = ocr_pdf(file) logger.debug("Extracting text...") text_pdf = read_pdf(file.filename, ocred_pdf) logger.debug("Saving to DB") create_file(file.read(), ocred_pdf, file.filename, text_pdf) # !! SECOND HASH return text_pdf
哈希函数代码
def md5_it(file): return hashlib.md5(file).hexdigest()
原因分析
核心问题是文件对象的指针位置偏移:
- 调用
file.read()后,文件指针会直接移动到文件末尾 - 后续再次调用
file.read()时,会返回空字节串b'' - 基于空内容计算的MD5哈希是固定值(
d41d8cd98f00b204e9800998ecf8427e),导致后续所有哈希计算都复用这个错误结果
解决步骤
有两种可靠的修复方式:
方式1:一次性读取文件内容到内存
先把文件完整内容读取到内存变量中,后续所有操作都基于这个变量,避免多次操作文件指针:
@app.route("/file_upload", methods=['POST']) def file_upload(): uploaded_files = request.files.getlist("pdf") for file in uploaded_files: logger.info(f"Processing item {file.filename}") # 一次性读取文件内容到内存 file_content = file.read() # 使用内存中的内容计算哈希 file_hash = md5_it(file_content) print(file_hash) in_db = get_file(file_hash) if in_db: logger.info(f"File already in the database, skipping...") print(in_db) text_pdf = in_db.text else: logger.debug("Doing OCR...") # 如果ocr_pdf需要读取文件内容,先重置指针 file.seek(0) ocred_pdf = ocr_pdf(file) logger.debug("Extracting text...") text_pdf = read_pdf(file.filename, ocred_pdf) logger.debug("Saving to DB") # 直接传入内存中的文件内容,无需再次read() create_file(file_content, ocred_pdf, file.filename, text_pdf) return text_pdf
方式2:每次读取后重置文件指针
如果必须多次读取文件对象,每次读取后调用file.seek(0)将指针重置到文件开头:
@app.route("/file_upload", methods=['POST']) def file_upload(): uploaded_files = request.files.getlist("pdf") for file in uploaded_files: logger.info(f"Processing item {file.filename}") # 第一次读取并计算哈希 file_hash = md5_it(file.read()) print(file_hash) # 重置指针到开头 file.seek(0) in_db = get_file(file_hash) if in_db: logger.info(f"File already in the database, skipping...") print(in_db) text_pdf = in_db.text else: logger.debug("Doing OCR...") ocred_pdf = ocr_pdf(file) logger.debug("Extracting text...") text_pdf = read_pdf(file.filename, ocred_pdf) logger.debug("Saving to DB") # 再次重置指针后读取文件内容 file.seek(0) create_file(file.read(), ocred_pdf, file.filename, text_pdf) return text_pdf
注意事项
- 如果
ocr_pdf函数内部会调用file.read(),必须确保调用前指针处于文件开头位置 - 优先选择方式1,减少重复IO操作,提升接口性能
内容的提问来源于stack exchange,提问作者Mike
相关产品推荐
相关产品推荐

