MediaProjection截图OCR识别异常,CPU占用影响论是否成立?求方案
Android两种截图方式的OCR识别差异及资源消耗疑问
问题背景
- Android平台采用两种截图方式:
- 执行命令:
/system/bin/screencap -p $path - 使用
MediaProjectionAPI
- 执行命令:
- 同一屏幕下用
Tesseract做OCR识别时,结果差异显著:- 用
/system/bin/screencap截图可得到预期识别结果 - 用
MediaProjectionAPI截图无法正确识别全部或部分文本,必须通过二值化算法预处理图像
- 用
截图实现细节
- 查阅
screencap源码,确认其采用PNG压缩、ARGB_8888像素格式、100%质量参数 - 我使用
MediaProjectionAPI生成Bitmap的代码如下:
public class ImageTransmogrifier implements ImageReader.OnImageAvailableListener { private final int width; private final int height; private final ImageReader imageReader; private final ScreenshotService svc; private Bitmap latestBitmap=null; ImageTransmogrifier(ScreenshotService svc) { this.svc=svc; Display display=svc.getWindowManager().getDefaultDisplay(); Point size=new Point(); display.getRealSize(size); int width=size.x; int height=size.y; while (width*height > (2<<19)) { width=width>>1; height=height>>1; } this.width=width; this.height=height; imageReader=ImageReader.newInstance(width, height, PixelFormat.RGBA_8888, 2); imageReader.setOnImageAvailableListener(this, svc.getHandler()); } @Override public void onImageAvailable(ImageReader reader) { final Image image=imageReader.acquireLatestImage(); if (image!=null) { Image.Plane[] planes=image.getPlanes(); ByteBuffer buffer=planes[0].getBuffer(); int pixelStride=planes[0].getPixelStride(); int rowStride=planes[0].getRowStride(); int rowPadding=rowStride - pixelStride * width; int bitmapWidth=width + rowPadding / pixelStride; if (latestBitmap == null || latestBitmap.getWidth() != bitmapWidth || latestBitmap.getHeight() != height) { if (latestBitmap != null) { latestBitmap.recycle(); } latestBitmap=Bitmap.createBitmap(bitmapWidth, height, Bitmap.Config.ARGB_8888); } latestBitmap.copyPixelsFromBuffer(buffer); image.close(); ByteArrayOutputStream baos=new ByteArrayOutputStream(); Bitmap cropped=Bitmap.createBitmap(latestBitmap, 0, 0, width, height); cropped.compress(Bitmap.CompressFormat.PNG, 100, baos); byte[] newPng=baos.toByteArray(); svc.processImage(newPng); } } Surface getSurface() { return(imageReader.getSurface()); } int getWidth() { return(width); } int getHeight() { return(height); } void close() { imageReader.close(); } }
疑问
有人提出:MediaProjection占用过多CPU资源,导致OCR可用算力不足,固定时间内识别精度下降;而screencap是单次截图,资源消耗远低于持续流传输,因此无需预处理。
请问:
- 该观点是否有依据?
- 若观点成立,应该更换替代方案还是仅对图像进行预处理?
内容的提问来源于stack exchange,提问作者zaxunobi
相关产品推荐
相关产品推荐

