You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C#调用文档库API递归获取全量内容超时问题求优化方案

文档库批量拉取性能优化方案

原方案瓶颈分析

  • 完全串行请求:所有API调用按顺序执行,每次只能处理1个文件夹,大量嵌套场景下总耗时等于所有单次API请求耗时之和,极易触发超时
  • 异步能力浪费:方法声明了async Task但未使用异步IO调用API,同步阻塞线程白白浪费性能
  • 无容错机制:单请求失败直接跳过对应文件夹,没有重试兜底,可能出现数据缺失

优化实现方案

1. 并行异步请求改造

使用SemaphoreSlim控制并发数,避免触发API限流,同时批量处理多个文件夹请求,相比串行耗时可降低数倍,示例代码如下:

// 需引入命名空间:System.Collections.Concurrent、System.Threading
public async Task<List<string>> RetrieveDocuments(string baseUrl, string rootId, int maxConcurrency = 10)
{
    var files = new ConcurrentBag<string>();
    var pendingFolderIds = new ConcurrentQueue<string>();
    pendingFolderIds.Enqueue(rootId);
    var semaphore = new SemaphoreSlim(maxConcurrency);
    var processingTasks = new List<Task>();

    while (pendingFolderIds.TryDequeue(out var currentFolderId))
    {
        await semaphore.WaitAsync();
        processingTasks.Add(Task.Run(async () =>
        {
            try
            {
                // 替换为实际异步API调用逻辑
                var response = await GetFolderContentAsync(baseUrl, currentFolderId);
                if (response == null) return;

                if (response.Folders != null)
                {
                    foreach (var subFolder in response.Folders)
                    {
                        pendingFolderIds.Enqueue(subFolder.FolderId);
                    }
                }

                if (response.Files != null)
                {
                    foreach (var file in response.Files)
                    {
                        files.Add(resourceRegex.Replace(file.Path, "/"));
                    }
                }
            }
            finally
            {
                semaphore.Release();
            }
        }));
    }

    await Task.WhenAll(processingTasks);
    return files.ToList();
}

// 示例异步API调用方法,可根据实际场景改造
private async Task<FolderResponse> GetFolderContentAsync(string baseUrl, string folderId)
{
    using var httpClient = new HttpClient();
    // 可按需添加请求头、超时配置、重试逻辑等
    httpClient.Timeout = TimeSpan.FromSeconds(10);
    return await httpClient.GetFromJsonAsync<FolderResponse>($"{baseUrl}/folder/{folderId}");
}

2. 额外优化点

  • 限流适配:根据文档库API的公开限流规则调整maxConcurrency参数,避免触发频率限制导致请求失败
  • 超时拆分:如果全量拉取总耗时仍超出阈值,可拆分任务先返回已拉取的文件列表,后续通过增量接口补全剩余数据
  • 缓存复用:重复拉取场景下可缓存已拉取过的文件夹内容,避免重复请求
  • 重试机制:给API调用增加指数退避重试逻辑,处理偶发的网络波动或限流拦截

提示:如果文档库API支持批量查询多个文件夹内容,优先使用批量接口替代单个请求,可进一步大幅降低总请求数和耗时。

内容的提问来源于stack exchange,提问作者Hershika Sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 16:39:03