You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP中array_merge与union性能差异的Zend底层原理探究

Great question—let’s dig into the Zend Engine internals to explain why array_merge() and the array union operator (+) have such staggering performance differences, especially in PHP 7. We’ll start with behavioral differences (since implementation follows behavior), then dive straight into the C-level code logic.

Core Behavioral Differences First

Before we get to the engine, let’s recap the key behavior that drives their performance:

  • Numeric keys: array_merge() resets numeric keys to sequential 0-based indexes, while + preserves the original numeric keys from the left array (ignoring any right-side elements with matching numeric keys).
  • String keys: array_merge() overwrites left-side string keys with right-side values, while + prioritizes left-side string keys entirely (right-side duplicates are ignored).

Zend Engine Implementation: array_merge()

The array_merge() function maps to the zend_array_merge C function in the Zend Engine. Here’s a simplified breakdown of its workflow:

  1. Precompute total size: It first iterates through all input arrays to calculate the total number of elements, then pre-allocates a new array with this size to minimize reallocations.
  2. Iterate and reindex every element: For each element in each input array:
    • It checks if the key is numeric. If yes, it discards the original key and uses an auto-incrementing counter to assign a new sequential index.
    • For string keys, it directly adds the element to the new array, overwriting any existing string key with the same name.
  3. Hidden overhead:
    • The key type check (numeric vs string) happens for every single element, adding per-element overhead.
    • Maintaining the auto-incrementing index for numeric keys requires extra state tracking.
    • Even with pre-allocation, edge cases like duplicate string keys can trigger expensive reallocations and memory copies.

In your benchmark, this per-element reindexing is the main culprit for the 7-second runtime: 20,000 elements mean 20,000 key checks and index assignments, which adds up quickly.

Zend Engine Implementation: Array Union (+ Operator)

The array union operator uses the zend_hash_merge C function with the ZEND_HASH_MERGE_LEFT flag. This is a far more efficient path:

  1. Copy the left array: First, it creates a shallow copy of the left-hand array (since PHP arrays follow value semantics).
  2. Batch merge with existence checks: It iterates through the right-hand array, but only adds elements where the key (numeric or string) does not already exist in the copied left array.
  3. No reindexing, minimal checks:
    • Numeric keys are preserved exactly as they are—no counter, no reindexing. The engine just checks if the key exists in the hash table (an O(1) operation thanks to Zend’s optimized hash table implementation).
    • String keys follow the same existence check; right-side duplicates are ignored entirely, aligning with the operator’s "left takes priority" behavior.
  4. Efficient memory handling: The copy of the left array uses Zend’s dynamic hash table expansion (doubling capacity when needed), which is optimized for speed. Since there’s no reindexing, there’s no extra memory churn from shifting elements.

In your benchmark, this is why the union operation takes "0 seconds": existence checks are fast, there’s no reindexing overhead, and the bulk merge is handled by low-level, optimized hash table code.

Why Your Benchmark Numbers Look So Extreme

Let’s tie this back to your results:

  • Runtime: array_merge()’s per-element reindexing is O(n) with high constant factors, while union’s bulk hash checks are O(n) with near-zero constant overhead. For 20,000 elements, this difference is massive.
  • Memory peaks: array_merge()’s higher peak memory comes from temporary state during reindexing and potential reallocations if the precomputed size is off. Union’s lower peak is because it starts with a copy of the left array and only adds missing elements, using Zend’s efficient hash table expansion.

Key Takeaways

  • Use + when you need to preserve numeric keys and only add elements that don’t already exist on the left—its performance is unbeatable for this use case.
  • Use array_merge() when you need sequential numeric indexes or want to overwrite string keys from left to right across multiple arrays.
  • The core performance gap boils down to key handling: array_merge() does heavy lifting to reindex numeric keys, while union just does fast hash table existence checks with no key modification.

内容的提问来源于stack exchange,提问作者Lukas Lukac

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:03:08