如何使用Python或VB.NET检测扫描PDF中的红色线条
扫描PDF红色线条检测实现方案
Python 实现
核心思路
扫描类PDF本质是图像集合,我们先将PDF每页转换为图像,再通过颜色识别提取红色区域,最后判断区域是否符合线条的形态(高长宽比)。
依赖安装
pip install pymupdf opencv-python numpy
代码示例
import fitz # PyMuPDF import cv2 import numpy as np def detect_red_lines_in_pdf(pdf_path, red_threshold=(0, 0, 100), line_aspect_ratio=5): # 打开PDF doc = fitz.open(pdf_path) red_line_pages = [] for page_num in range(len(doc)): page = doc.load_page(page_num) # 转换为高分辨率图片(DPI=300) pix = page.get_pixmap(dpi=300) # 转换为OpenCV格式的BGR图像 img = np.frombuffer(pix.samples, dtype=np.uint8).reshape(pix.height, pix.width, 3) img = cv2.cvtColor(img, cv2.COLOR_RGB2BGR) # 提取红色区域(RGB阈值,可根据实际调整) lower_red = np.array([red_threshold[0], red_threshold[1], red_threshold[2]]) upper_red = np.array([255, 255, 255]) mask = cv2.inRange(img, lower_red, upper_red) # 查找轮廓 contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) has_red_line = False for cnt in contours: x, y, w, h = cv2.boundingRect(cnt) # 判断是否为线条:长宽比大于设定值,且面积过滤零散像素 if max(w, h) / min(w, h) > line_aspect_ratio and cv2.contourArea(cnt) > 20: has_red_line = True break if has_red_line: red_line_pages.append(page_num + 1) # 页码从1开始计数 doc.close() return red_line_pages # 使用示例 if __name__ == "__main__": result = detect_red_lines_in_pdf("scan_file.pdf") if result: print(f"检测到红色线条的页码:{result}") else: print("未检测到红色线条")
关键说明
- 红色阈值:可根据扫描件实际红色深浅调整
red_threshold参数 - 线条判断:通过轮廓的长宽比和面积过滤零散红色像素,避免误判
- DPI设置:转图片时设置高DPI(如300)能提升细线条的检测精度
VB.NET 实现
核心思路
利用PdfiumViewer将PDF页转换为Bitmap,遍历像素识别红色区域,再通过连续红色像素的长度判断是否为线条。
依赖安装
通过NuGet安装以下包:
PdfiumViewer(PDF转图像)System.Drawing.Common(图像处理)
代码示例
Imports PdfiumViewer Imports System.Drawing Imports System.Collections.Generic Public Class RedLineDetector Public Shared Function DetectRedLines(pdfPath As String, redThreshold As Integer, minLineLength As Integer) As List(Of Integer) Dim redLinePages As New List(Of Integer)() Using document As PdfDocument = PdfDocument.Load(pdfPath) For pageNum As Integer = 0 To document.PageCount - 1 Using bitmap As Bitmap = document.Render(pageNum, 300, 300, PdfRenderFlags.Annotations) Dim hasRedLine As Boolean = False Dim currentLineLength As Integer = 0 For y As Integer = 0 To bitmap.Height - 1 For x As Integer = 0 To bitmap.Width - 1 Dim pixel As Color = bitmap.GetPixel(x, y) ' 判断是否为红色:R值远大于G、B值,且超过阈值 If pixel.R > redThreshold AndAlso pixel.G < pixel.R * 0.3 AndAlso pixel.B < pixel.R * 0.3 Then currentLineLength += 1 If currentLineLength >= minLineLength Then hasRedLine = True Exit For End If Else currentLineLength = 0 End If Next If hasRedLine Then Exit For Next If hasRedLine Then redLinePages.Add(pageNum + 1) ' 页码从1开始 End If End Using Next End Using Return redLinePages End Function ' 使用示例 Public Shared Sub Main() Dim result = DetectRedLines("scan_file.pdf", 180, 20) If result.Count > 0 Then Console.WriteLine($"检测到红色线条的页码:{String.Join(", ", result)}") Else Console.WriteLine("未检测到红色线条") End If End Sub End Class
关键说明
- 红色判断:通过R、G、B的比例识别红色,避免受亮度影响
- 线条长度:通过连续红色像素的累计长度过滤零散红点
- 资源释放:使用
Using语句确保PDF和Bitmap资源被正确释放
内容的提问来源于stack exchange,提问作者Paavak V
相关产品推荐
相关产品推荐

