TensorFlow 1.14中Eager与Graph模式性能异常问题咨询
Great question—it’s counterintuitive since we usually expect Graph Mode to outperform Eager Mode for repeated runs, but your results make total sense once we dig into TensorFlow 1.x’s execution quirks and your test setup. Let’s break down the key reasons:
1. Session.run() and feed_dict Overhead Dominates Small Computations
Your Graph Mode code relies on a tf.placeholder and feeds the numpy array val via feed_dict every time you call sess.run(). This introduces two major overheads that completely overshadow the Graph’s optimization benefits for tiny models:
- Python-to-C++ Boundary Crossings: Each
sess.run()requires jumping between Python’s runtime and TensorFlow’s C++ backend, adding fixed overhead per call. For a simple Dense layer, this overhead makes up most of the total runtime. - Repeated Data Copying: Feeding a numpy array via
feed_dictmeans copying data from Python’s memory into TensorFlow’s C++ memory pool on every run. Your Eager Mode code uses atf.Tensordirectly (created withtf.math.scalar_mul), so there’s no back-and-forth data copying between domains.
2. Warmup and Optimization Disparities
While you added warmup runs for both modes, TF1.x’s Graph Mode doesn’t fully optimize the graph on the first invocation in all cases. Meanwhile, Eager Mode in TF1.14 already applies lightweight just-in-time optimizations for tensor operations—for very small computations, these optimizations make Eager feel snappier than Graph Mode, which carries the extra baggage of session management.
3. tf.function Eliminates Graph Mode’s Legacy Overheads
When you wrap your Eager code with tf.function, you’re using AutoGraph to convert the Python function into an optimized graph—but unlike your manual Graph Mode setup, tf.function in TF1.14 avoids feed_dict entirely. It uses TensorFlow’s native tensor handles directly, so there’s no repeated data copying or Python/C++ boundary crossing for each call. This amplifies the performance gap because your manual Graph Mode code is tied to the old Session API’s limitations.
How to Adjust the Benchmark to See Graph Mode’s Strengths
To get a fairer comparison that shows Graph Mode’s intended advantages, modify your Graph Mode code to eliminate feed_dict overhead:
- Use a
tf.Variableto store your input tensor (so data stays in TensorFlow’s memory):import tensorflow as tf import numpy as np import timeit graph = tf.Graph() with graph.as_default(): val = tf.Variable(np.ones((10, 200, 200, 3)) * 3.5, dtype=tf.float32) f = tf.keras.layers.Flatten()(val) c = tf.keras.layers.Dense(512, activation='relu')(f) with tf.Session(graph=graph) as sess: sess.run(tf.global_variables_initializer()) # Warmup sess.run(c) print(timeit.timeit(lambda: sess.run(c), number=1000)/1000) - Alternatively, batch multiple runs into a single
sess.run()call to minimize boundary crossings.
Key Takeaway
TensorFlow 1.x’s Graph Mode shines for large, complex computations that run many times without repeated data feeding—think training loops with big batches, distributed training, or inference serving. For tiny models with repeated small invocations, the overhead of the Session API and feed_dict easily outweighs Graph Mode’s optimization benefits. TF2.x fixed this by unifying Eager and Graph modes via tf.function, which eliminates most of these legacy overheads by default.
内容的提问来源于stack exchange,提问作者Pelups

