You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS环境下OpenCV Python脚本因资源超限被终止,求优化方案

问题

我有一段Python/OpenCV代码,功能是将填写完成的文档(new.png)与参考文档(ref.png)对齐,并将结果保存至output.png。但在运行时偶尔会在bash中出现「Killed」提示,推测是资源耗尽导致,可能代码存在效率问题,希望得到优化建议。

运行环境:Python 3.5.3、OpenCV 4.4.0、Linux内核版本4.14.301-224.520.amzn2.x86_64。

代码如下:

import sys
import cv2
import numpy as np

if len(sys.argv) != 4:
  print('USAGE')
  print('  python3 diff.py ref.png new.png output.png')
  sys.exit()

GOOD_MATCH_PERCENT = 0.15

def alignImages(im1, im2):
  # Convert images to grayscale
  #im1Gray = cv2.cvtColor(im1, cv2.COLOR_BGR2GRAY)
  #im2Gray = cv2.cvtColor(im2, cv2.COLOR_BGR2GRAY)

  # Detect ORB features and compute descriptors.
  orb = cv2.AKAZE_create()
  #orb = cv2.ORB_create(500)
  keypoints1, descriptors1 = orb.detectAndCompute(im1, None)
  keypoints2, descriptors2 = orb.detectAndCompute(im2, None)

  # Match features.
  matcher = cv2.DescriptorMatcher_create(cv2.DESCRIPTOR_MATCHER_BRUTEFORCE_HAMMING)
  matches = matcher.match(descriptors1, descriptors2, None)

  # Sort matches by score
  matches.sort(key=lambda x: x.distance, reverse=False)

  # Remove not so good matches
  numGoodMatches = int(len(matches) * GOOD_MATCH_PERCENT)
  print(matches[numGoodMatches].distance)
  matches = matches[:numGoodMatches]

  # Draw top matches
  imMatches = cv2.drawMatches(im1, keypoints1, im2, keypoints2, matches, None)
  cv2.imwrite("matches.jpg", imMatches)

  # Extract location of good matches
  points1 = np.zeros((len(matches), 2), dtype=np.float32)
  points2 = np.zeros((len(matches), 2), dtype=np.float32)

  for i, match in enumerate(matches):
    points1[i, :] = keypoints1[match.queryIdx].pt
    points2[i, :] = keypoints2[match.trainIdx].pt

  # Find homography
  h, mask = cv2.findHomography(points1, points2, cv2.RANSAC)

  # Use homography
  height, width = im2.shape
  im1Reg = cv2.warpPerspective(im1, h, (width, height))

  return im1Reg, h

def removeOverlap(refBW, newBW):
  # invert each
  refBW = 255 - refBW
  newBW = 255 - newBW

  # get absdiff
  xor = cv2.absdiff(refBW, newBW)

  result = cv2.bitwise_and(xor, newBW)

  # invert
  result = 255 - result

  return result

def offset(img, xOffset, yOffset):
  # The number of pixels
  num_rows, num_cols = img.shape[:2]

  # Creating a translation matrix
  translation_matrix = np.float32([ [1,0,xOffset], [0,1,yOffset] ])

  # Image translation
  img_translation = cv2.warpAffine(img, translation_matrix, (num_cols,num_rows), borderValue = (255,255,255))

  return img_translation


# the ink will often bleed out on printouts ever so slightly
# to eliminate that we'll apply a "jitter" of sorts

refFilename = sys.argv[1]
imFilename = sys.argv[2]
outFilename = sys.argv[3]

imRef = cv2.imread(refFilename, cv2.IMREAD_GRAYSCALE)
im = cv2.imread(imFilename, cv2.IMREAD_GRAYSCALE)

imNew, h = alignImages(im, imRef)

refBW = cv2.threshold(imRef, 0, 255, cv2.THRESH_BINARY+cv2.THRESH_OTSU)[1]
newBW = cv2.threshold(imNew, 0, 255, cv2.THRESH_BINARY+cv2.THRESH_OTSU)[1]

for x in range (-2, 2):
  for y in range (-2, 2):
    newBW = removeOverlap(offset(refBW, x, y), newBW)

cv2.imwrite(outFilename, newBW)
优化建议
  • 替换特征检测器,减少特征数量:AKAZE计算成本远高于ORB,且默认会生成大量特征。直接启用注释中的cv2.ORB_create(500),限制特征数量为500,大幅降低内存占用和计算时间。
  • 移除特征匹配可视化代码:drawMatches和cv2.imwrite("matches.jpg", imMatches)会生成大尺寸匹配图,对核心功能无帮助,直接删除这部分代码,节省内存和IO时间。
  • 向量化特征点提取,替代Python循环:原代码用for循环逐个赋值points1和points2,效率极低。改用numpy向量化操作:
    points1 = np.array([keypoints1[m.queryIdx].pt for m in matches], dtype=np.float32)
    points2 = np.array([keypoints2[m.trainIdx].pt for m in matches], dtype=np.float32)
    
    避免Python层面的循环开销,提升运行速度。
  • 用形态学操作替代嵌套偏移循环:原代码中16次循环处理墨水扩散的逻辑,内存和时间成本高。改用OpenCV内置的形态学开运算替代,利用底层优化实现:
    # 替换原有嵌套循环
    kernel = np.ones((2,2), np.uint8)
    newBW = cv2.morphologyEx(newBW, cv2.MORPH_OPEN, kernel)
    
    同样能消除墨水扩散噪点,速度提升显著。
  • 缩小图像做特征匹配,再还原对齐:如果文档尺寸过大,先将imRef和im缩小到合适尺寸(比如宽高减半)进行特征检测和匹配,得到单应性矩阵后,再用原图执行warpPerspective。大幅减少特征检测阶段的内存占用。
  • 手动释放无用变量:在alignImages函数中,keypoints1、descriptors1等变量在计算完单应性矩阵后不再使用,赋值为None帮助垃圾回收释放内存:
    keypoints1 = descriptors1 = keypoints2 = descriptors2 = matches = None
    

内容的提问来源于stack exchange,提问作者neubert

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 21:50:15