You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Delphi Berlin 10.1中Parallel.For比普通for循环耗时更长求助

嘿,我来帮你分析下为啥OtlParallel的Parallel.For反而比普通循环慢,以及怎么解决这个问题~

为什么Parallel.For反而更慢?

首先得明确:并行循环不是“万能加速药”,它本身存在线程创建、调度、同步的开销。如果你的循环体执行逻辑本身很轻量(比如你代码里的字符串查找、拼接),而且slSMP的元素数量不是特别大,那并行带来的开销会直接盖过并行的收益,导致整体耗时更长。

另外还有个隐藏的致命问题:你在并行循环里直接调用slTempSearchList.Add,但TStringList不是线程安全的!多个线程同时操作同一个TStringList会引发数据竞争,轻则结果错乱,重则程序崩溃,这也是你必须先解决的问题。

优化方案

1. 先解决线程安全,再优化性能

并行场景下不能直接共享TStringList,正确的做法是让每个线程先收集自己的结果,最后再合并到主列表,这样既安全又能减少同步开销。

示例代码:

var
  TempLists: array of TStringList;
  I, ThreadIdx: Integer;
begin
  // 为每个线程初始化临时列表
  SetLength(TempLists, OtlParallel.TaskConfig.MaxThreadCount);
  for I := 0 to High(TempLists) do
    TempLists[I] := TStringList.Create;

  try
    OtlParallel.Parallel.For(0, slSMP.Count - 1)
      .Execute(procedure(const AIndex: Integer)
        var
          CurrentList: TStringList;
          IsHit: Boolean;
          NameStr, ValueStr, FileNameStr: string;
        begin
          // 获取当前线程对应的临时列表
          ThreadIdx := GetCurrentThreadId mod Length(TempLists);
          CurrentList := TempLists[ThreadIdx];
          
          // 提前缓存属性值,减少重复访问开销
          NameStr := slSMP.Names[AIndex];
          ValueStr := slSMP.ValueFromIndex[AIndex];
          
          // 提前计算是否命中,避免重复调用字符串处理函数
          FileNameStr := ExtractFileName(ExcludeTrailingPathDelimiter(NameStr));
          IsHit := System.StrUtils.ContainsText(ValueStr, ThisSearchTerm) 
                 or System.StrUtils.ContainsText(FileNameStr, ThisSearchTerm);
          
          if IsHit then
            CurrentList.Add(NameStr + ValueStr);
        end);

    // 合并所有临时列表到主列表
    slTempSearchList.BeginUpdate;
    try
      for I := 0 to High(TempLists) do
        slTempSearchList.AddStrings(TempLists[I]);
    finally
      slTempSearchList.EndUpdate;
    end;
  finally
    // 释放临时列表
    for I := 0 to High(TempLists) do
      TempLists[I].Free;
  end;
end;

2. 调整并行粒度,减少调度开销

如果slSMP的元素数量不算特别多(比如几千条以内),可以把元素分成大块处理,而不是每个元素单独开一个任务——这样能大幅减少任务调度的开销。

示例:

OtlParallel.Parallel.For(0, (slSMP.Count div 100) - 1) // 每100个元素为一块
  .Execute(procedure(const ABlockIdx: Integer)
    var
      StartIdx, EndIdx, I: Integer;
      BlockList: TStringList;
      IsHit: Boolean;
      NameStr, ValueStr, FileNameStr: string;
    begin
      StartIdx := ABlockIdx * 100;
      EndIdx := Min(StartIdx + 99, slSMP.Count - 1);
      BlockList := TStringList.Create;
      try
        for I := StartIdx to EndIdx do
        begin
          NameStr := slSMP.Names[I];
          ValueStr := slSMP.ValueFromIndex[I];
          FileNameStr := ExtractFileName(ExcludeTrailingPathDelimiter(NameStr));
          IsHit := System.StrUtils.ContainsText(ValueStr, ThisSearchTerm) 
                 or System.StrUtils.ContainsText(FileNameStr, ThisSearchTerm);
          if IsHit then
            BlockList.Add(NameStr + ValueStr);
        end;
        // 加锁合并到主列表
        TMonitor.Enter(slTempSearchList);
        try
          slTempSearchList.AddStrings(BlockList);
        finally
          TMonitor.Exit(slTempSearchList);
        end;
      finally
        BlockList.Free;
      end;
    end);

3. 先优化串行代码,再考虑并行

很多时候,优化后的串行代码效率已经足够,甚至比并行更快。你可以先对原串行代码做这些优化:

  • 提前缓存ThisSearchTerm的大小写转换结果,避免ContainsText重复转换字符串
  • 用BeginUpdate/EndUpdate包裹Add操作,减少TStringList的内存重分配和UI重绘开销
  • 把属性访问结果存到局部变量,减少重复访问的开销

优化后的串行代码:

var
  I: Integer;
  LowerSearchTerm: string;
  NameStr, ValueStr, FileNameStr: string;
begin
  LowerSearchTerm := LowerCase(ThisSearchTerm);
  slTempSearchList.BeginUpdate;
  try
    for I := 0 to slSMP.Count - 1 do
    begin
      NameStr := slSMP.Names[I];
      ValueStr := slSMP.ValueFromIndex[I];
      FileNameStr := LowerCase(ExtractFileName(ExcludeTrailingPathDelimiter(NameStr)));
      // 用Pos代替ContainsText,手动控制大小写转换,提升效率
      if (Pos(LowerSearchTerm, LowerCase(ValueStr)) > 0) 
         or (Pos(LowerSearchTerm, FileNameStr) > 0) then
      begin
        slTempSearchList.Add(NameStr + ValueStr);
      end;
    end;
  finally
    slTempSearchList.EndUpdate;
  end;
end;

4. 手动指定并行线程数

默认情况下OtlParallel会使用和CPU核心数相同的线程数,但线程数过多会导致上下文切换频繁,反而降低效率。你可以根据自己的CPU配置手动调整:

OtlParallel.Parallel.For(0, slSMP.Count - 1)
  .MaxDegreeOfParallelism(4) // 比如4核CPU就设4,根据实际情况调整
  .Execute(...);
总结

优先优化串行代码的效率,解决线程安全问题,再根据数据量调整并行的粒度和线程数。如果数据量不大,优化后的串行代码可能已经足够快,没必要强行用并行。

内容的提问来源于stack exchange,提问作者user1580348

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:11:36