You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速Django应用中受数据库调用限制的大规模哈希比较?

汉明距离哈希匹配的性能优化问题

需求背景

我正在开发一个Django Web应用,需要对超过1000万个64字符长度的字符串哈希进行比较,最终返回相似度最高的100个哈希(相似度基于汉明距离,即相同字符的数量)。

基本流程:用户发起POST请求→views.py处理并生成输入哈希→传入以下核心处理代码:

from hexhamming import hamming_distance_string
from django.conf import settings
import psycopg2 as pg
import pandas as pd

def _hamming_distance(hash1, hash2):
    return hamming_distance_string(hash1, hash2)

def compare_image_hashes(hash):

    result_df = pd.DataFrame({'key': [], 'hash': [], 'difference':[]})  # Empty dataframe for all results

    MAX_CURSOR_PULL = settings.MAX_CURSOR_PULL   # Max amount of results for the PG cursor
    SIMILARITY = settings.SIMILARITY # Value to determine if an image is similar
    LIMIT = settings.LIMIT


    conn = pg.connect(
            database=settings.DATABASE['DATABASE_NAME'],
            user=settings.DATABASE['DATABASE_USER'],
            password=settings.DATABASE['DATABASE_PASSWORD'],
            host=settings.DATABASE['DATABASE_HOST'],
            port=settings.DATABASE['DATABASE_PORT']
        )

    cursor = conn.cursor('hash-cursor')
    cursor.itersize = MAX_CURSOR_PULL # Redudant limit

    query = "SELECT hash, key FROM hashes LIMIT %s;" % LIMIT
    cursor.execute(query)

    chunk = cursor.fetchmany(MAX_CURSOR_PULL)
    columns = [desc[0] for desc in cursor.description]

    while chunk:    # Continue while there are still results
        df = pd.DataFrame(chunk, columns=columns)

        df['difference'] = [_hamming_distance(hash, x.strip()) for x in df['hash']]    # Add a column for the hash difference of the image

        if result_df.empty: # If there are no results
            result_df = df.query('difference <= %s' % SIMILARITY)
        else:   # If there are results concat the two dataframes
            result_df = pd.concat([result_df, df.query('difference <= %s' % SIMILARITY)])

        chunk = cursor.fetchmany(MAX_CURSOR_PULL)


    result_df = result_df.sort_values('difference', kind="mergesort")[:100]

    return result_df

当前性能问题

  • 响应时间长达约90秒,偶尔出现超时错误
  • 部署环境:AWS EC2(30vCPU、256GB内存、3TB磁盘),该实例同时托管另外两个应用;哈希存储在AWS RDS Postgres实例中
  • 技术栈:Python 3.10.13、Pandas 2.1.X、Psycopg2 2.X、Django 3.12.4
  • 核心瓶颈:需要遍历1000多万条哈希数据,数据库调用频次高,EC2无法高效批量获取并处理全量数据

已尝试/考虑的方案

  • 使用hexhamming库加速汉明距离计算,虽有性能提升但仍未解决速度/超时问题
  • 在本地尝试并行化哈希获取与计算逻辑,无明显改善
  • 初步优化思路:
    • 将哈希获取与计算逻辑迁移到独立服务器或Lambda,避免影响其他Web应用;或在内存中缓存数据库静态副本,减少PG调用次数
    • 采用服务器外的并行化方式,分块处理哈希的获取与汉明距离计算

寻求帮助

需要可行的性能优化方案,涵盖架构调整、代码优化、数据库层面改进等方向。

内容的提问来源于stack exchange,提问作者Jesse Fentress

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 20:02:17