You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Kotlin实现无Google API、无弹窗的后台语音识别

实现无弹窗、不依赖Google服务的Android离线语音识别

你的现有代码依赖系统自带的RecognizerIntent,这会唤起第三方语音识别Activity(通常是Google的),所以必然有弹窗且依赖Google服务。要实现需求,必须改用本地离线语音识别库,完全在App内部处理语音流。下面以Vosk为例(开源、支持多语言、离线运行),给出完整修改方案:

步骤1:添加依赖与配置

在项目的build.gradle(Module level)中添加以下内容:

repositories {
    mavenCentral()
}

dependencies {
    implementation 'org.vosk:vosk-android:0.3.45'
}

同时在AndroidManifest.xml中添加录音权限:

<uses-permission android:name="android.permission.RECORD_AUDIO" />

步骤2:准备语言模型

下载对应语言的离线模型(支持en-US、fil-PH),解压后放在App的assets目录下,比如:

  • assets/model-en-us/(英文模型)
  • assets/model-fil-ph/(他加禄语模型)

步骤3:替换原有识别逻辑

删除原来基于RecognizerIntent的代码,改用Vosk的API实现后台识别:

import android.Manifest
import android.content.pm.PackageManager
import android.os.Bundle
import android.view.View
import android.widget.AdapterView
import androidx.appcompat.app.AppCompatActivity
import androidx.core.app.ActivityCompat
import androidx.core.content.ContextCompat
import org.vosk.LibVosk
import org.vosk.Model
import org.vosk.Recognizer
import org.vosk.android.RecognitionListener
import org.vosk.android.SpeechService
import org.json.JSONObject
import java.io.IOException

class MainActivity : AppCompatActivity(), AdapterView.OnItemSelectedListener {
    private val RECORD_AUDIO_PERMISSION_REQUEST = 100
    private var currentLangCode = "en-US"
    private val langNames = arrayOf("English", "Tagalog")
    private val langCodes = arrayOf("en-US", "fil-PH")
    private val modelPaths = arrayOf("model-en-us", "model-fil-ph")
    
    private var model: Model? = null
    private var speechService: SpeechService? = null
    private var isListening = false

    override fun onCreate(savedInstanceState: Bundle?) {
        super.onCreate(savedInstanceState)
        setContentView(R.layout.your_layout) // 替换成你的布局ID

        // 初始化Vosk库
        LibVosk.setLogLevel(LibVosk.LOG_LEVEL_WARN)

        // 检查录音权限
        if (ContextCompat.checkSelfPermission(this, Manifest.permission.RECORD_AUDIO)
            != PackageManager.PERMISSION_GRANTED) {
            ActivityCompat.requestPermissions(
                this,
                arrayOf(Manifest.permission.RECORD_AUDIO),
                RECORD_AUDIO_PERMISSION_REQUEST
            )
        }

        // 语言选择Spinner监听(保留原逻辑)
        yourSpinner?.onItemSelectedListener = this // 替换成你的Spinner ID

        // 麦克风按钮点击事件
        micbtn?.setOnClickListener {
            if (isListening) {
                stopRecognition()
                micbtn?.setText("Start Listening") // 替换成你的按钮文本逻辑
            } else {
                if (ContextCompat.checkSelfPermission(this, Manifest.permission.RECORD_AUDIO)
                    == PackageManager.PERMISSION_GRANTED) {
                    startRecognition(currentLangCode)
                    micbtn?.setText("Stop Listening")
                } else {
                    ActivityCompat.requestPermissions(
                        this,
                        arrayOf(Manifest.permission.RECORD_AUDIO),
                        RECORD_AUDIO_PERMISSION_REQUEST
                    )
                }
            }
        }
    }

    private fun loadModel(langCode: String) {
        val modelPath = when(langCode) {
            "en-US" -> modelPaths[0]
            "fil-PH" -> modelPaths[1]
            else -> modelPaths[0]
        }
        try {
            model?.close() // 关闭旧模型
            model = Model(assets, modelPath)
        } catch (e: IOException) {
            e.printStackTrace()
            // 处理模型加载失败的情况,比如提示用户
        }
    }

    private fun startRecognition(langCode: String) {
        if (model == null || model?.modelPath != modelPaths[langCodes.indexOf(langCode)]) {
            loadModel(langCode)
        }
        try {
            val recognizer = Recognizer(model, 16000.0f)
            speechService = SpeechService(recognizer, 16000.0f)
            speechService?.setRecognitionListener(object : RecognitionListener {
                override fun onResult(hypothesis: String?) {
                    // 识别结果回调,处理最终文本
                    val result = JSONObject(hypothesis).getString("text")
                    if (edtext?.text?.isNotEmpty() == true) {
                        edtext?.append(" ")
                    }
                    edtext?.append(result)
                }

                override fun onPartialResult(hypothesis: String?) {
                    // 可以处理实时部分识别结果,不需要的话留空
                }

                override fun onFinalResult(hypothesis: String?) {
                    isListening = false
                    micbtn?.setText("Start Listening")
                }

                override fun onError(e: Exception?) {
                    isListening = false
                    micbtn?.setText("Start Listening")
                    e?.printStackTrace()
                }

                override fun onTimeout() {
                    isListening = false
                    micbtn?.setText("Start Listening")
                }
            })
            speechService?.startListening()
            isListening = true
        } catch (e: IOException) {
            e.printStackTrace()
        }
    }

    private fun stopRecognition() {
        speechService?.stop()
        speechService?.shutdown()
        isListening = false
    }

    override fun onRequestPermissionsResult(
        requestCode: Int,
        permissions: Array<out String>,
        grantResults: IntArray
    ) {
        super.onRequestPermissionsResult(requestCode, permissions, grantResults)
        if (requestCode == RECORD_AUDIO_PERMISSION_REQUEST) {
            if (grantResults.isNotEmpty() && grantResults[0] == PackageManager.PERMISSION_GRANTED) {
                startRecognition(currentLangCode)
                micbtn?.setText("Stop Listening")
            }
        }
    }

    override fun onItemSelected(adapterView: AdapterView<*>?, view: View?, i: Int, l: Long) {
        currentLangCode = langCodes[i]
        // 如果正在识别,可切换模型后重启识别,或提示用户先停止
        if (isListening) {
            stopRecognition()
            // 可选:自动切换语言后重新开始
            // startRecognition(currentLangCode)
            // micbtn?.setText("Stop Listening")
        }
    }

    override fun onNothingSelected(adapterView: AdapterView<*>?) {
        // 无操作
    }

    override fun onDestroy() {
        super.onDestroy()
        speechService?.shutdown()
        model?.close()
    }
}

关键说明

  • 完全移除了对RecognizerIntent的依赖,所有识别逻辑在App内部完成,无弹窗
  • 使用Vosk离线模型,不依赖Google服务,支持后台运行
  • 保留了原有的语言切换逻辑,可在English和Tagalog之间切换
  • 处理了录音权限的动态申请,符合Android权限规范
  • 识别过程中可以实时停止,点击按钮切换状态

注意:语言模型文件较大(英文模型约1.8G,他加禄语模型约1.5G),建议引导用户在WiFi环境下载,或者将模型放在外部存储,避免Apk体积过大。

内容的提问来源于stack exchange,提问作者Yuno

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 20:42:18