You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将pytesseract提取的字符串或列表转为Python DataFrame?

将OCR提取的文本转换为Pandas DataFrame

问题背景

你编写了一段基于pytesseract和cv2的屏幕截图文本提取代码,当前得到两种格式的输出:

  • 单条字符串:包含表头与空格分隔的行数据,存在部分识别错误
  • 多次提取存储的字符串列表
    需要将这两种格式的输出转换为Pandas DataFrame。

原代码

import pyautogui
import cv2
import pytesseract
import numpy as np
from PIL import ImageGrab, Image, ImageEnhance
import time
from pytesseract import Output


size = 1000, 1000
chats = []
pytesseract.pytesseract.tesseract_cmd = 'My Patch to Tesseract'

def ScreenSearch():
 while True:
    time.sleep(1)
    img = (ImageGrab.grab(bbox =(1290, 247, 1586, 517)))
    img = img.resize(size, Image.ANTIALIAS)

    gray = cv2.cvtColor(np.array(img), cv2.COLOR_BGR2GRAY)

    # Perform text extraction
    data = pytesseract.image_to_string(gray, lang='eng', config='--psm 6')

    print(data)
    print(type(data))
    
    chats.append(data)
    print(chats)
    print(type(chats))
    
ScreenSearch()

输出示例

Preco(USDT) Quantia(BTC) Total
23624.21 0.00617 145.76138
23624.04 0.02000 472.48080
23624.00 0.00100 YRS tt)
23623.60 0.00650 153.55340
23623.37 0.00842 198.90878
23623.36 0.00846 199.85363
23623.28 0.00636 150.24406
23623.27 0.01913 451.91316
23623.01 0.00640 151.18726
23622.98 0.00675 159.45512
23622.85 0.00052 12.28388
23622.84 0.00210 49.60796

<class 'str'>
['Preco(USDT) Quantia(BTC) Total\n23624.21 0.00617 145.76138\n23624.04 0.02000 472.48080\n23624.00 0.00100 YRS tt)\n23623.60 0.00650 153.55340\n23623.37 0.00842 198.90878\n23623.36 0.00846 199.85363\n23623.28 0.00636 150.24406\n23623.27 0.01913 451.91316\n23623.01 0.00640 151.18726\n23622.98 0.00675 159.45512\n23622.85 0.00052 12.28388\n23622.84 0.00210 49.60796\n']
<class 'list'>

解决方案

首先导入pandas库,然后分别处理两种格式的输出:

1. 处理单条提取的字符串

将字符串按换行符分割,按空格拆分每行数据,构造DataFrame时跳过格式异常的行,并尝试转换数值类型:

import pandas as pd

def str_to_df(raw_str):
    # 按换行分割行并过滤空行
    lines = [line.strip() for line in raw_str.split('\n') if line.strip()]
    if not lines:
        return pd.DataFrame()
    
    # 提取表头
    headers = lines[0].split()
    data_rows = []
    
    # 处理数据行,跳过列数不匹配的异常行
    for line in lines[1:]:
        parts = line.split()
        if len(parts) == len(headers):
            data_rows.append(parts)
        else:
            print(f"跳过格式异常的行: {line}")
    
    # 构造DataFrame并尝试转换数值类型
    df = pd.DataFrame(data_rows, columns=headers)
    for col in headers:
        try:
            df[col] = df[col].astype(float)
        except ValueError:
            pass  # 转换失败则保留字符串类型
    return df

2. 处理字符串列表

遍历列表中的每个字符串,转换为DataFrame后合并并去重:

def list_to_df(str_list):
    dfs = []
    for s in str_list:
        df = str_to_df(s)
        if not df.empty:
            dfs.append(df)
    
    if not dfs:
        return pd.DataFrame()
    
    # 合并所有DataFrame并去重
    merged_df = pd.concat(dfs, ignore_index=True).drop_duplicates()
    return merged_df

集成到原代码中

修改原代码,在提取文本后直接转换为DataFrame:

import pyautogui
import cv2
import pytesseract
import numpy as np
from PIL import ImageGrab, Image, ImageEnhance
import time
import pandas as pd

size = 1000, 1000
chats = []
pytesseract.pytesseract.tesseract_cmd = 'My Patch to Tesseract'

def str_to_df(raw_str):
    lines = [line.strip() for line in raw_str.split('\n') if line.strip()]
    if not lines:
        return pd.DataFrame()
    
    headers = lines[0].split()
    data_rows = []
    for line in lines[1:]:
        parts = line.split()
        if len(parts) == len(headers):
            data_rows.append(parts)
        else:
            print(f"跳过格式异常的行: {line}")
    
    df = pd.DataFrame(data_rows, columns=headers)
    for col in headers:
        try:
            df[col] = df[col].astype(float)
        except ValueError:
            pass
    return df

def list_to_df(str_list):
    dfs = []
    for s in str_list:
        df = str_to_df(s)
        if not df.empty:
            dfs.append(df)
    
    if not dfs:
        return pd.DataFrame()
    
    merged_df = pd.concat(dfs, ignore_index=True).drop_duplicates()
    return merged_df

def ScreenSearch():
 while True:
    time.sleep(1)
    img = ImageGrab.grab(bbox=(1290, 247, 1586, 517))
    img = img.resize(size, Image.ANTIALIAS)

    gray = cv2.cvtColor(np.array(img), cv2.COLOR_BGR2GRAY)

    # 提取文本
    data = pytesseract.image_to_string(gray, lang='eng', config='--psm 6')

    # 转换单条字符串为DataFrame
    single_df = str_to_df(data)
    print("单条转换结果:\n", single_df)
    
    chats.append(data)
    # 转换列表为合并后的DataFrame
    list_df = list_to_df(chats)
    print("列表合并结果:\n", list_df)
    
ScreenSearch()

额外优化建议

  • 图像预处理:增加cv2.threshold阈值处理,提高文本清晰度,减少OCR识别错误
  • 异常记录:将识别异常的行写入日志文件,方便后续排查问题
  • 重复数据:如果多次截图提取到重复内容,drop_duplicates()可有效去重

内容的提问来源于stack exchange,提问作者Eric Alves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 23:15:53