You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Cloud Run上EF Core连接Cloud SQL报错:无法创建Socket的排查与解决

问题描述

我在Google Cloud Run上托管基于.NET Core 8的应用,采用Google Cloud Functions Framework,Entity Framework版本为8.0.7,PostgreSQL数据库部署在Google Cloud SQL,且已配置Cloud Run的Cloud SQL连接。

连接字符串配置如下:

var connectionString = new NpgsqlConnectionStringBuilder
{
    SslMode = SslMode.Disable, // 即使sslmode设为disable,Cloud SQL Auth Proxy仍会提供加密连接
    Host = $"/cloudsql/{environmentVariableService.DatabaseConnectionName}",
    Username = environmentVariableService.DatabaseUser,
    Password = environmentVariableService.DatabasePassword,
    Database = environmentVariableService.DatabaseName,
    Pooling = true,
    ApplicationName = environmentVariableService.FunctionTarget,
    MinPoolSize = 60,
    CommandTimeout = 30,
    Port = 5432
};

services.AddDbContext<ApplicationDbContext>(
    options =>
        options.UseNpgsql(connectionString.ConnectionString), ServiceLifetime.Transient);

使用注入的ApplicationDbContext访问数据库时,每天会多次出现错误,大部分时间正常,报错发生在以下代码处:

public DocumentInfoRepository(ApplicationDbContext dbContext, RequestCancellationBase requestCancellationBase)
{
    _dbContext = dbContext;
    _requestCancellationBase = requestCancellationBase;
}

public async Task<DBModel.DocumentInfo> GetDocumentById(Guid id)
{
    // 错误发生在这一行 ⬇️
    return await _dbContext.DocumentInfos.SingleOrDefaultAsync(e => e.Id == id, _requestCancellationBase);
}

报错信息:

could not create socket for "db_connection": listen unix /tmp/cloudsql-proxy-tmp/db_connection/.s.PGSQL.5432: bind: invalid argument

    An unexpected error occured.
    System.InvalidOperationException: An exception has been raised that is likely due to a transient failure.
     ---&gt; Npgsql.NpgsqlException (0x80004005): Failed to connect to /cloudsql/environment-production:northamerica-northeast1:postgres-database-name/.s.PGSQL.5432
     ---&gt; System.Net.Sockets.SocketException (99): Cannot assign requested address
       ...(省略后续堆栈信息)

目前只能通过重新部署Cloud Run容器暂时解决问题,请问错误原因是什么?如何彻底解决?


错误原因分析
  1. 连接池配置过载:设置的MinPoolSize = 60远高于Cloud Run单实例的资源承载能力。Cloud Run实例的套接字、文件描述符数量有限,强制保留大量闲置连接会快速耗尽实例资源,导致无法创建新连接。
  2. Unix套接字连接未回收:通过Cloud SQL Auth Proxy建立的Unix套接字连接,若连接池长期持有未使用的连接,加上MinPoolSize的强制保留规则,会造成套接字资源泄漏,触发Cannot assign requested address错误。
  3. DbContext生命周期不合理:将ApplicationDbContext注册为Transient生命周期,会导致每次请求创建全新的DbContext实例,连接池无法高效复用已有连接,进一步加剧资源消耗。

彻底解决步骤

1. 调整连接池参数

降低MinPoolSize至合理范围,同时配置闲置连接超时,避免资源长期占用:

var connectionString = new NpgsqlConnectionStringBuilder
{
    // 保留原有其他配置
    MinPoolSize = 5, // 建议设置为5-10,适配Cloud Run实例资源
    MaxPoolSize = 50, // 根据实例并发量调整,默认100可适当降低
    IdleTimeout = 180 // 闲置连接3分钟后自动回收,减少资源占用
};

2. 优化DbContext生命周期

将ApplicationDbContext改为Scoped生命周期(这是AddDbContext的默认配置),让每个请求复用一个DbContext实例,提升连接池复用效率:

services.AddDbContext<ApplicationDbContext>(
    options => options.UseNpgsql(connectionString.ConnectionString));
// 无需显式指定ServiceLifetime.Scoped,默认即可

3. 启用EF重试机制

通过Entity Framework的重试策略处理 transient 错误,自动重试无效连接请求:

services.AddDbContext<ApplicationDbContext>(options =>
    options.UseNpgsql(connectionString.ConnectionString)
           .EnableRetryOnFailure(
               maxRetryCount: 5,
               maxRetryDelay: TimeSpan.FromSeconds(10),
               errorCodesToAdd: null));

4. 调整Cloud Run实例资源

若实例频繁出现资源耗尽,可适当提升实例的CPU/内存配置(例如从1CPU/2GB调整至2CPU/4GB),增加实例可使用的套接字和文件描述符数量。

5. 优化Cloud Run缩容策略

在Cloud Run控制台配置合理的缩容冷却时间,让闲置实例及时缩容,释放持有的连接池资源,避免闲置实例占用过多资源。


内容的提问来源于stack exchange,提问作者Frédéric Fect

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 02:14:54