You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求提供Google Cloud OCR图文识别完整分步实操示例(PHP/JS)

完整Google Cloud Vision OCR图文识别实操指南(PHP/JS)

别担心,我来帮你把整个流程捋明白,下面是针对PHP和JavaScript的完整可运行示例,包括库的获取、引入以及代码细节的解释:


PHP 实现步骤

1. 安装Google Cloud Vision PHP库

Google Cloud的PHP SDK依赖Composer管理,先确保你已经安装了Composer,然后在项目根目录执行以下命令:

composer require google/cloud-vision

安装完成后,Composer会自动生成vendor/autoload.php,这是自动加载所有依赖库的入口文件。

2. 代码示例与关键点解释

下面是完整的OCR识别代码,我会逐一解释你疑惑的部分:

<?php
// 引入Composer自动加载器,必须通过这一步加载所有Google Cloud依赖库
require __DIR__ . '/vendor/autoload.php';

// 你提到的`use Google\Cloud\Vision\V1\ImageAnnotatorClient;`是PHP的命名空间引入
// 作用是让你可以直接用`ImageAnnotatorClient`短类名,不用每次写冗长的完整类路径
use Google\Cloud\Vision\V1\ImageAnnotatorClient;
use Google\Cloud\Vision\V1\Image;
use Google\Cloud\Vision\V1\Feature;
use Google\Cloud\Vision\V1\Feature\Type;

// 替换成你实际的服务账户密钥文件路径
$keyFilePath = __DIR__ . '/service-account-key.json';

try {
    // 初始化OCR客户端,指定密钥文件
    $imageAnnotator = new ImageAnnotatorClient([
        'credentials' => $keyFilePath
    ]);

    // 替换成你要识别的本地图片路径
    $imagePath = __DIR__ . '/test-image.jpg';
    // 读取图片二进制内容
    $imageContent = file_get_contents($imagePath);
    $image = (new Image())->setContent($imageContent);

    // 设置任务类型为文本识别
    $feature = (new Feature())->setType(Type::TEXT_DETECTION);
    // 发送识别请求
    $response = $imageAnnotator->annotateImage([
        'image' => $image,
        'features' => [$feature]
    ]);

    // 提取识别结果
    $textAnnotations = $response->getTextAnnotations();
    if (!empty($textAnnotations)) {
        // 第一个结果是完整的识别文本,后续是逐行/逐词的细分结果
        echo "识别到的完整文本:\n";
        echo $textAnnotations[0]->getDescription() . "\n\n";

        echo "逐段识别结果:\n";
        foreach ($textAnnotations as $index => $annotation) {
            if ($index === 0) continue; // 跳过第一个完整文本
            echo "- " . $annotation->getDescription() . "\n";
        }
    } else {
        echo "未识别到任何文本";
    }

    // 关闭客户端释放资源
    $imageAnnotator->close();
} catch (Exception $e) {
    echo "识别出错: " . $e->getMessage();
}
?>

关于你疑惑的命名空间:

  • namespace Google\Cloud\Samples\Vision;是官方示例代码的命名空间,你自己写代码时完全可以不用这个命名空间,直接写在全局命名空间下即可(就像上面的示例一样)。
  • use Google\Cloud\Vision\V1\ImageAnnotatorClient;是PHP的标准用法,用来简化类名引用,避免每次都写冗长的完整类路径。

JavaScript(Node.js)实现步骤

1. 安装Google Cloud Vision Node.js库

确保你已经安装了Node.js,然后在项目根目录执行:

npm install @google-cloud/vision

2. 代码示例

const vision = require('@google-cloud/vision');
const fs = require('fs');

// 初始化OCR客户端,指定服务账户密钥文件路径
const client = new vision.ImageAnnotatorClient({
    keyFilename: './service-account-key.json' // 替换成你的密钥文件路径
});

async function detectText() {
    try {
        // 替换成你要识别的本地图片路径
        const imagePath = './test-image.jpg';
        // 读取图片内容
        const imageBuffer = fs.readFileSync(imagePath);
        // 发送OCR请求
        const [result] = await client.textDetection(imageBuffer);
        const detections = result.textAnnotations;

        if (detections.length > 0) {
            console.log('识别到的完整文本:');
            console.log(detections[0].description + '\n');

            console.log('逐段识别结果:');
            detections.slice(1).forEach(text => {
                console.log('- ' + text.description);
            });
        } else {
            console.log('未识别到任何文本');
        }
    } catch (err) {
        console.error('识别出错:', err);
    }
}

// 执行识别函数
detectText();

浏览器端JS说明

如果需要在浏览器中调用Vision API,强烈建议通过你的后端服务中转(避免密钥暴露在前端),或者使用Google的OAuth 2.0授权流程(不推荐直接在前端使用服务账户密钥)。


关键注意事项

  • 确保你的服务账户已经被授予Cloud Vision API Editor或Cloud Vision API Viewer权限(在GCP控制台的IAM页面配置)。
  • 支持的图片格式包括JPG、PNG、GIF、BMP等,单张图片大小不超过20MB。
  • 如果图片存储在Google Cloud Storage中,可以直接传入GCS路径(比如gs://your-bucket/test-image.jpg),无需读取本地文件。

内容的提问来源于stack exchange,提问作者Europeuser

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:43:27