You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Sign-In令牌刷新超时,如何制定有效重试策略?

Optimal Retry Strategy for Google ID Token Refresh Failures (silentSignIn Timeouts)

I need to refresh Google tokens used for server login validation, which have a TTL of about 1 hour. Most of the time the refresh process works fine, but sometimes calling Google to get a new token fails for various reasons. I'm looking for a viable retry mechanism for Google token refresh.

Here's the existing code in my app:

private val googleSignInClient: GoogleSignInClient by lazy { 
    // This takes a measurable amount of time to compute, so do it lazily 
    val gso = GoogleSignInOptions.Builder(GoogleSignInOptions.DEFAULT_SIGN_IN)
        .requestIdToken(WEB_CLIENT_ID) // need this to get user ID token later
        .requestEmail()
        .build()
    GoogleSignIn.getClient(appContext, gso)
}

override fun getToken() = getRefreshedGoogleInfo()?.googleToken()

/**
 * Here we "silently sign in" to get a refreshed Google ID Token.
 *
 * This method might block, so do not call it from the main thread.
 */
private fun getRefreshedGoogleInfo(): GoogleUserInfo? {
    val task = googleSignInClient.silentSignIn()
    // If the task is already complete, return the result immediately
    if (task.isComplete) {
        val info = task.result.toGoogleUserInfo()
        Logger.v(TAG, "silentSignIn result from already-completed task = %s", info.toString())
        return info
    }
    // If the task is not complete, await up to 5s and return result, or null
    return try {
        val info = task.await().toGoogleUserInfo()
        Logger.v(TAG, "silentSignIn result from await task = %s", info.toString())
        info
    } catch (e: Exception) {
        Logger.e(TAG, e, "silentSignIn result from await task = null\nerror = ${e.localizedMessage}")
        null
    }
}

private fun Task<GoogleSignInAccount>.await() = Tasks.await(this, 5, TimeUnit.SECONDS)

Sometimes task.await() fails due to timeout. I tried retrying immediately a few times, but subsequent retries after the first failure are always ineffective. The Google official docs don't provide effective guidance for this scenario. What's the optimal retry strategy?


Answer

Great question—this is a super common pain point with token refresh flows, and the core issue with your immediate retry approach is that you’re not accounting for why the failure happened or giving the system time to bounce back. Here’s a battle-tested strategy tailored specifically to Google’s silentSignIn flow:

1. First, Filter Which Failures Are Actually Retryable

Not every failure is worth retrying—you need to skip retries for permanent issues that won’t fix themselves:

  • Retry only transient errors: Focus on timeouts, network-related IOExceptions, and Google’s ApiException with status codes like CommonStatusCodes.NETWORK_ERROR or CommonStatusCodes.SERVICE_UNAVAILABLE. These are temporary glitches that might resolve with a little time.
  • Skip permanent failures: Don’t waste cycles on ApiException with codes like CommonStatusCodes.SIGN_IN_REQUIRED (user needs to re-authenticate manually), CommonStatusCodes.INVALID_ACCOUNT, or permission errors. Retrying these will just give you the same result.

2. Use Exponential Backoff for Retries

Immediate retries almost never work because transient issues (like network blips or Google server throttling) need a small window to recover. Exponential backoff means you wait longer between each retry:

  • Start with a short 1-second delay after the first failure
  • Double the delay each subsequent attempt (2s → 4s → 8s)
  • Cap the maximum delay at 10 seconds to avoid waiting unreasonably long

This is the standard recommended approach for cloud services (including Google) because it reduces server load and drastically boosts your chance of success.

3. Limit Total Retry Attempts

Set a hard limit (e.g., 3 retries) to prevent your code from getting stuck in an infinite loop if the issue persists. Once you hit the limit, fall back to a user-facing recovery step (like prompting them to re-sign in manually).

Modified Code Implementation

Here’s how to update your getRefreshedGoogleInfo method to include this strategy:

private fun getRefreshedGoogleInfo(): GoogleUserInfo? {
    val maxRetries = 3
    var currentRetry = 0
    var delayMs = 1000L // Start with 1 second delay

    while (currentRetry <= maxRetries) {
        val task = googleSignInClient.silentSignIn()
        try {
            val info = if (task.isComplete) {
                task.result.toGoogleUserInfo()
            } else {
                task.await().toGoogleUserInfo()
            }
            Logger.v(TAG, "silentSignIn succeeded on attempt ${currentRetry + 1}: %s", info.toString())
            return info
        } catch (e: Exception) {
            currentRetry++
            // Check if this failure is retryable
            val isRetryable = when (e) {
                is TimeoutException -> true
                is IOException -> true
                is ApiException -> {
                    when (e.statusCode) {
                        CommonStatusCodes.NETWORK_ERROR,
                        CommonStatusCodes.SERVICE_UNAVAILABLE -> true
                        else -> false
                    }
                }
                else -> false
            }

            if (!isRetryable || currentRetry > maxRetries) {
                Logger.e(TAG, e, "silentSignIn failed permanently after ${currentRetry} attempts: ${e.localizedMessage}")
                return null
            }

            // Wait before retrying
            Logger.w(TAG, "silentSignIn failed on attempt ${currentRetry}, retrying in ${delayMs}ms: ${e.localizedMessage}")
            Thread.sleep(delayMs)
            // Double the delay for next retry, cap at 10s
            delayMs = minOf(delayMs * 2, 10000L)
        }
    }
    return null
}

Why This Works Better Than Immediate Retries

Google’s silentSignIn might cache transient failure states briefly, or your device could still be stuck in a bad network state. By adding a growing delay, you give both your device and Google’s servers time to recover—something immediate retries don’t account for. This will make your refresh flow far more reliable.

内容的提问来源于stack exchange,提问作者TonyR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:30:32