You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中基于图像特征而非坐标为肿瘤样本图像添加标注

自动给肿瘤样本图像标注蛋白质表达数值的方案

R 实现方案

核心思路是用OCR识别图像中的区域编号,获取其位置后将对应数值标注到编号旁,依赖tesseract做文字识别、magick处理图像。

步骤1:安装并加载依赖包

install.packages(c("tesseract", "magick", "dplyr", "stringr", "tidyr"))
library(tesseract)
library(magick)
library(dplyr)
library(stringr)
library(tidyr)

步骤2:整理标注数据

把宽格式数据转成长格式,提取区域编号(比如从Patient1_001中提取001):

# 示例数据
expr_data <- data.frame(
  Ki_67 = c(0.0162, 0.0707, 0.177),
  row.names = c("Patient1_001", "Patient1_002", "Patient1_003")
) %>% t() %>% as.data.frame() %>% rownames_to_column("Protein") %>% pivot_longer(-Protein, names_to = "Sample", values_to = "Value")

# 提取3位数字的区域编号
expr_data <- expr_data %>% mutate(Region_ID = str_extract(Sample, "\\d{3}"))

步骤3:单张图像标注

# 读取目标图像
img <- image_read("Patient1.png")

# OCR识别图像中的文字及坐标信息
ocr_result <- ocr_data(img)

# 筛选出3位数字的区域编号
region_ocr <- ocr_result %>% filter(str_detect(word, "\\d{3}"))

# 匹配数值并标注
for (i in 1:nrow(region_ocr)) {
  id <- region_ocr$word[i]
  value <- expr_data$Value[expr_data$Region_ID == id]
  
  if (length(value) > 0) {
    # 在编号右侧偏移10像素处标注数值,避免重叠
    x <- region_ocr$x[i] + region_ocr$width[i] + 10
    y <- region_ocr$y[i] + region_ocr$height[i]/2
    
    img <- image_annotate(img, as.character(value), 
                          size = 12, color = "red",
                          location = geometry_point(x, y))
  }
}

# 保存标注后的图像
image_write(img, "Patient1_annotated.png")

步骤4:批量处理

遍历文件夹下所有PNG/TIFF文件,自动完成标注:

img_dir <- "./tumor_images/"
img_files <- list.files(img_dir, pattern = "\\.(png|tiff)$", full.names = TRUE)

for (file in img_files) {
  # 提取样本名称(比如从Patient1.png中获取Patient1)
  sample_name <- str_extract(basename(file), "^[^.]+")
  img <- image_read(file)
  ocr_result <- ocr_data(img)
  region_ocr <- ocr_result %>% filter(str_detect(word, "\\d{3}"))
  
  for (i in 1:nrow(region_ocr)) {
    id <- region_ocr$word[i]
    target_sample <- paste0(sample_name, "_", id)
    value <- expr_data$Value[expr_data$Sample == target_sample]
    
    if (length(value) > 0) {
      x <- region_ocr$x[i] + region_ocr$width[i] + 10
      y <- region_ocr$y[i] + region_ocr$height[i]/2
      img <- image_annotate(img, as.character(value), size = 12, color = "red", location = geometry_point(x, y))
    }
  }
  
  # 保存标注后的图像
  image_write(img, file.path(img_dir, paste0(sample_name, "_annotated.png")))
}

Python 实现方案

用pytesseract做OCR识别,Pillow处理图像,pandas整理数据,逻辑和R一致,图像预处理更灵活。

步骤1:安装依赖

pip install pytesseract pillow pandas opencv-python

注:Windows用户需提前安装Tesseract OCR引擎,并在代码中指定引擎路径

步骤2:整理标注数据

import pandas as pd

# 示例数据
expr_data = pd.DataFrame({
    "Patient1_001": [0.0162],
    "Patient1_002": [0.0707],
    "Patient1_003": [0.177]
}, index=["Ki-67"])

# 转成长格式并提取区域编号
expr_data = expr_data.T.reset_index()
expr_data.columns = ["Sample", "Ki-67"]
expr_data["Region_ID"] = expr_data["Sample"].str.extract(r"(\d{3})")

步骤3:单张图像标注

from PIL import Image, ImageDraw, ImageFont
import pytesseract

# Windows用户需指定Tesseract路径
# pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'

# 读取图像
img = Image.open("Patient1.png")
draw = ImageDraw.Draw(img)

# OCR识别文字及坐标,返回字典格式结果
ocr_dict = pytesseract.image_to_data(img, output_type=pytesseract.Output.DICT)

# 遍历识别结果,筛选3位数字的区域编号
n_boxes = len(ocr_dict['text'])
for i in range(n_boxes):
    text = ocr_dict['text'][i].strip()
    if text.isdigit() and len(text) == 3:
        # 获取编号的坐标信息
        x, y, w, h = ocr_dict['left'][i], ocr_dict['top'][i], ocr_dict['width'][i], ocr_dict['height'][i]
        # 匹配对应数值
        value = expr_data.loc[expr_data["Region_ID"] == text, "Ki-67"].values
        if len(value) > 0:
            # 在编号右侧标注数值,偏移10像素
            annotate_x = x + w + 10
            annotate_y = y + h // 2
            # 设置字体(优先系统字体, fallback到默认字体)
            try:
                font = ImageFont.truetype("arial.ttf", 12)
            except:
                font = ImageFont.load_default()
            draw.text((annotate_x, annotate_y), str(value[0]), fill="red", font=font)

# 保存标注后的图像
img.save("Patient1_annotated.png")

步骤4:批量处理

import os

img_dir = "./tumor_images/"
img_files = [f for f in os.listdir(img_dir) if f.lower().endswith(('.png', '.tiff'))]

for file in img_files:
    file_path = os.path.join(img_dir, file)
    sample_name = os.path.splitext(file)[0]
    img = Image.open(file_path)
    draw = ImageDraw.Draw(img)
    ocr_dict = pytesseract.image_to_data(img, output_type=pytesseract.Output.DICT)
    
    n_boxes = len(ocr_dict['text'])
    for i in range(n_boxes):
        text = ocr_dict['text'][i].strip()
        if text.isdigit() and len(text) == 3:
            x, y, w, h = ocr_dict['left'][i], ocr_dict['top'][i], ocr_dict['width'][i], ocr_dict['height'][i]
            target_sample = f"{sample_name}_{text}"
            if target_sample in expr_data["Sample"].values:
                value = expr_data.loc[expr_data["Sample"] == target_sample, "Ki-67"].values[0]
                annotate_x = x + w + 10
                annotate_y = y + h // 2
                try:
                    font = ImageFont.truetype("arial.ttf", 12)
                except:
                    font = ImageFont.load_default()
                draw.text((annotate_x, annotate_y), str(value), fill="red", font=font)
    
    # 保存标注后的图像
    img.save(os.path.join(img_dir, f"{sample_name}_annotated.png"))

注意事项

  • OCR识别准确率依赖图像质量,若原图像编号模糊,可先做预处理(灰度化、二值化),R的magick和Python的opencv都能实现
  • 可根据需求调整标注的字体大小、颜色、位置偏移量
  • 批量处理前建议先用少量图像测试,确认匹配逻辑和标注位置正确后再扩大范围

内容的提问来源于stack exchange,提问作者mfeldbauer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 02:37:09