如何比普通循环更快处理大型对象数组?递归实现性能疑问
如何让对象数组的批量操作比普通循环更快?
我想知道怎么让整个对象数组的操作速度比普通循环更快。
首先定义了一个对单个元素执行操作的模板函数:
template <class T> inline void operate(T &_this) { /* do something with _this */ }
我最初实现的循环版本代码如下:
template <class T> void loop(const size_t _total, T *_array) { for(unsigned _cursor(0); _cursor < _total; _cursor++) { T *_t(_array + _cursor); ::operate<T>(*_t); } }
后来我打算改成线程友好的实现,于是把原loop()改成了递归的iterate()函数,同时新增了apply()函数——通过分片处理数组,以此实现线程中断和调用栈增长的管理,修改后的代码如下:
template <class T> inline void iterate(const size_t &_total, T *_these, void (*_that)(T &)) { if(_these != 0) { if(_total > 1) ::iterate<T>(_total - 1, _these + 1, _that); _that(*_these); } } template <class T, size_t SLICE = 1024> inline void apply(const size_t &_total, T *_these, void (*_that)(T &)) { if(_total > 0) { const bool _recurse(_total > SLICE); if(_recurse == true) ::apply<T>(SLICE, _total - SLICE, _these + SLICE, _that); ::iterate<T>((_recurse == true) ? SLICE : _total, _these, _that); } }
目前代码能正常运行,但我认为凭借apply()和iterate()中的递归调用逻辑,整个数组的处理速度应该会比原loop()版本有大幅提升。我想确认:Borland C++ Builder 5编译器能否完成这类优化——比如内联operate()函数,同时对每个数组元素的参数做作用域处理?或者我是不是遗漏了某些关键点?这种优化思路到底是否正确?
内容的提问来源于stack exchange,提问作者blueperfect
相关产品推荐
相关产品推荐

