如何加速Django应用中受数据库调用限制的大规模哈希比较?
汉明距离哈希匹配的性能优化问题
需求背景
我正在开发一个Django Web应用,需要对超过1000万个64字符长度的字符串哈希进行比较,最终返回相似度最高的100个哈希(相似度基于汉明距离,即相同字符的数量)。
基本流程:用户发起POST请求→views.py处理并生成输入哈希→传入以下核心处理代码:
from hexhamming import hamming_distance_string from django.conf import settings import psycopg2 as pg import pandas as pd def _hamming_distance(hash1, hash2): return hamming_distance_string(hash1, hash2) def compare_image_hashes(hash): result_df = pd.DataFrame({'key': [], 'hash': [], 'difference':[]}) # Empty dataframe for all results MAX_CURSOR_PULL = settings.MAX_CURSOR_PULL # Max amount of results for the PG cursor SIMILARITY = settings.SIMILARITY # Value to determine if an image is similar LIMIT = settings.LIMIT conn = pg.connect( database=settings.DATABASE['DATABASE_NAME'], user=settings.DATABASE['DATABASE_USER'], password=settings.DATABASE['DATABASE_PASSWORD'], host=settings.DATABASE['DATABASE_HOST'], port=settings.DATABASE['DATABASE_PORT'] ) cursor = conn.cursor('hash-cursor') cursor.itersize = MAX_CURSOR_PULL # Redudant limit query = "SELECT hash, key FROM hashes LIMIT %s;" % LIMIT cursor.execute(query) chunk = cursor.fetchmany(MAX_CURSOR_PULL) columns = [desc[0] for desc in cursor.description] while chunk: # Continue while there are still results df = pd.DataFrame(chunk, columns=columns) df['difference'] = [_hamming_distance(hash, x.strip()) for x in df['hash']] # Add a column for the hash difference of the image if result_df.empty: # If there are no results result_df = df.query('difference <= %s' % SIMILARITY) else: # If there are results concat the two dataframes result_df = pd.concat([result_df, df.query('difference <= %s' % SIMILARITY)]) chunk = cursor.fetchmany(MAX_CURSOR_PULL) result_df = result_df.sort_values('difference', kind="mergesort")[:100] return result_df
当前性能问题
- 响应时间长达约90秒,偶尔出现超时错误
- 部署环境:AWS EC2(30vCPU、256GB内存、3TB磁盘),该实例同时托管另外两个应用;哈希存储在AWS RDS Postgres实例中
- 技术栈:Python 3.10.13、Pandas 2.1.X、Psycopg2 2.X、Django 3.12.4
- 核心瓶颈:需要遍历1000多万条哈希数据,数据库调用频次高,EC2无法高效批量获取并处理全量数据
已尝试/考虑的方案
- 使用
hexhamming库加速汉明距离计算,虽有性能提升但仍未解决速度/超时问题 - 在本地尝试并行化哈希获取与计算逻辑,无明显改善
- 初步优化思路:
- 将哈希获取与计算逻辑迁移到独立服务器或Lambda,避免影响其他Web应用;或在内存中缓存数据库静态副本,减少PG调用次数
- 采用服务器外的并行化方式,分块处理哈希的获取与汉明距离计算
寻求帮助
需要可行的性能优化方案,涵盖架构调整、代码优化、数据库层面改进等方向。
内容的提问来源于stack exchange,提问作者Jesse Fentress
相关产品推荐
相关产品推荐

