Delphi Berlin 10.1中Parallel.For比普通for循环耗时更长求助
嘿,我来帮你分析下为啥OtlParallel的Parallel.For反而比普通循环慢,以及怎么解决这个问题~
为什么Parallel.For反而更慢?
首先得明确:并行循环不是“万能加速药”,它本身存在线程创建、调度、同步的开销。如果你的循环体执行逻辑本身很轻量(比如你代码里的字符串查找、拼接),而且slSMP的元素数量不是特别大,那并行带来的开销会直接盖过并行的收益,导致整体耗时更长。
另外还有个隐藏的致命问题:你在并行循环里直接调用slTempSearchList.Add,但TStringList不是线程安全的!多个线程同时操作同一个TStringList会引发数据竞争,轻则结果错乱,重则程序崩溃,这也是你必须先解决的问题。
优化方案
1. 先解决线程安全,再优化性能
并行场景下不能直接共享TStringList,正确的做法是让每个线程先收集自己的结果,最后再合并到主列表,这样既安全又能减少同步开销。
示例代码:
var TempLists: array of TStringList; I, ThreadIdx: Integer; begin // 为每个线程初始化临时列表 SetLength(TempLists, OtlParallel.TaskConfig.MaxThreadCount); for I := 0 to High(TempLists) do TempLists[I] := TStringList.Create; try OtlParallel.Parallel.For(0, slSMP.Count - 1) .Execute(procedure(const AIndex: Integer) var CurrentList: TStringList; IsHit: Boolean; NameStr, ValueStr, FileNameStr: string; begin // 获取当前线程对应的临时列表 ThreadIdx := GetCurrentThreadId mod Length(TempLists); CurrentList := TempLists[ThreadIdx]; // 提前缓存属性值,减少重复访问开销 NameStr := slSMP.Names[AIndex]; ValueStr := slSMP.ValueFromIndex[AIndex]; // 提前计算是否命中,避免重复调用字符串处理函数 FileNameStr := ExtractFileName(ExcludeTrailingPathDelimiter(NameStr)); IsHit := System.StrUtils.ContainsText(ValueStr, ThisSearchTerm) or System.StrUtils.ContainsText(FileNameStr, ThisSearchTerm); if IsHit then CurrentList.Add(NameStr + ValueStr); end); // 合并所有临时列表到主列表 slTempSearchList.BeginUpdate; try for I := 0 to High(TempLists) do slTempSearchList.AddStrings(TempLists[I]); finally slTempSearchList.EndUpdate; end; finally // 释放临时列表 for I := 0 to High(TempLists) do TempLists[I].Free; end; end;
2. 调整并行粒度,减少调度开销
如果slSMP的元素数量不算特别多(比如几千条以内),可以把元素分成大块处理,而不是每个元素单独开一个任务——这样能大幅减少任务调度的开销。
示例:
OtlParallel.Parallel.For(0, (slSMP.Count div 100) - 1) // 每100个元素为一块 .Execute(procedure(const ABlockIdx: Integer) var StartIdx, EndIdx, I: Integer; BlockList: TStringList; IsHit: Boolean; NameStr, ValueStr, FileNameStr: string; begin StartIdx := ABlockIdx * 100; EndIdx := Min(StartIdx + 99, slSMP.Count - 1); BlockList := TStringList.Create; try for I := StartIdx to EndIdx do begin NameStr := slSMP.Names[I]; ValueStr := slSMP.ValueFromIndex[I]; FileNameStr := ExtractFileName(ExcludeTrailingPathDelimiter(NameStr)); IsHit := System.StrUtils.ContainsText(ValueStr, ThisSearchTerm) or System.StrUtils.ContainsText(FileNameStr, ThisSearchTerm); if IsHit then BlockList.Add(NameStr + ValueStr); end; // 加锁合并到主列表 TMonitor.Enter(slTempSearchList); try slTempSearchList.AddStrings(BlockList); finally TMonitor.Exit(slTempSearchList); end; finally BlockList.Free; end; end);
3. 先优化串行代码,再考虑并行
很多时候,优化后的串行代码效率已经足够,甚至比并行更快。你可以先对原串行代码做这些优化:
- 提前缓存
ThisSearchTerm的大小写转换结果,避免ContainsText重复转换字符串 - 用
BeginUpdate/EndUpdate包裹Add操作,减少TStringList的内存重分配和UI重绘开销 - 把属性访问结果存到局部变量,减少重复访问的开销
优化后的串行代码:
var I: Integer; LowerSearchTerm: string; NameStr, ValueStr, FileNameStr: string; begin LowerSearchTerm := LowerCase(ThisSearchTerm); slTempSearchList.BeginUpdate; try for I := 0 to slSMP.Count - 1 do begin NameStr := slSMP.Names[I]; ValueStr := slSMP.ValueFromIndex[I]; FileNameStr := LowerCase(ExtractFileName(ExcludeTrailingPathDelimiter(NameStr))); // 用Pos代替ContainsText,手动控制大小写转换,提升效率 if (Pos(LowerSearchTerm, LowerCase(ValueStr)) > 0) or (Pos(LowerSearchTerm, FileNameStr) > 0) then begin slTempSearchList.Add(NameStr + ValueStr); end; end; finally slTempSearchList.EndUpdate; end; end;
4. 手动指定并行线程数
默认情况下OtlParallel会使用和CPU核心数相同的线程数,但线程数过多会导致上下文切换频繁,反而降低效率。你可以根据自己的CPU配置手动调整:
OtlParallel.Parallel.For(0, slSMP.Count - 1) .MaxDegreeOfParallelism(4) // 比如4核CPU就设4,根据实际情况调整 .Execute(...);
总结
优先优化串行代码的效率,解决线程安全问题,再根据数据量调整并行的粒度和线程数。如果数据量不大,优化后的串行代码可能已经足够快,没必要强行用并行。
内容的提问来源于stack exchange,提问作者user1580348
相关产品推荐
相关产品推荐

