有序结构并行流reduce操作,identity不合规时是否与串行流输出一致?
Question 1: Do parallel streams derived from lists always behave the same as their serial counterparts, producing predictable identical output?
Short answer: No, they don’t guarantee identical output by default. The consistency between parallel and serial streams depends entirely on whether the operations you’re using meet the core requirements laid out in the Stream API specification:
- Stateless operations: Operations that don’t rely on or modify external state (e.g., a pure
mapfunction is safe;forEachthat updates an external counter is not). - Associative operations: For reduction operations like
reduce, the accumulator and combiner functions must be associative—meaning the order of combining results doesn’t change the final outcome. - Non-interfering operations: The stream source isn’t modified while the stream is being processed.
If any of these conditions are violated, parallel streams can produce wildly different results than their serial equivalents. For example, using a stateful filter or a non-associative accumulator will break consistency between parallel and serial runs.
Question 2: Can we assume parallel reduce on ordered structures (like lists) will match serial output, even when the identity violates the spec?
Your test runs 100 times and always returns true, but this is not a guarantee you can depend on long-term. Here’s the breakdown:
First, let’s revisit the spec requirement for the identity value in reduce: it must satisfy combiner.apply(identity, u) == u for every element u in the stream. In your code, the identity is "x", and combiner.apply("x", "A") returns "xA"—which is not equal to "A". This explicitly violates the API’s rules.
The consistent results you’re seeing right now are due to implementation details of the current JDK’s parallel stream splitting logic for ordered collections. Right now, the JVM splits your list into substreams, runs the reduce on each, then combines the results in a way that lines up with the serial stream’s concatenation order. But this isn’t part of the official spec—it’s just how the current version happens to work.
The critical point here: the Stream API labels results from non-compliant reduce calls as undefined behavior. There’s no guarantee that future JDK versions, or even alternative JVM implementations, will use the same splitting strategy. A small change in how parallel streams split ordered lists could lead to a result like "xAxEIxOxU" or another variation that doesn’t match the serial output.
Key Takeaway
Never rely on undefined behavior in the Stream API, even if it seems to work consistently today. To ensure parallel and serial streams produce identical results, you must strictly follow all the API’s rules: use a valid identity, ensure your accumulator and combiner are associative and compatible, and only use stateless, non-interfering operations.
内容的提问来源于stack exchange,提问作者Treefish Zhang

