Spark技术问题:函数内持久化DataFrame后仍重复执行操作
Hey there! Let's figure out why your DataFrame isn't staying cached between calls to getdf2() and fix this issue. The most common culprits boil down to how function-scoped variables work and missing caching logic—here's a breakdown of the key issues and solutions:
1. Local Variables Don't Persist Between Function Calls
When you define a DataFrame inside getdf2(), it's a local variable that gets created fresh every time you call the function. Once the function finishes running, Python cleans up those local variables, so there's nothing left to reuse on the next call.
Fix: Cache the DataFrame as a Function Attribute
You can store the cached DataFrame directly on the function itself (functions are objects in Python, so they can have attributes!). This way, the data sticks around between calls:
import pandas as pd def getdf2(): # Check if we already have a cached DataFrame if hasattr(getdf2, "_cached_df"): return getdf2._cached_df # Your original time-consuming operations here print("Running expensive DataFrame creation...") # For testing df = pd.DataFrame({"col1": range(10000), "col2": range(10000, 20000)}) # Save the DataFrame to the function's attribute getdf2._cached_df = df return df
Now, the first call to getdf2() will run the expensive code, but every subsequent call will return the cached DataFrame immediately.
2. Missing Explicit Caching Logic
If you're relying on implicit persistence (hoping the DataFrame stays in memory without extra code), that won't work—you need to explicitly tell Python to cache the result.
Fix: Use functools.cache (Python 3.9+)
For a cleaner approach, use the built-in cache decorator to automatically cache the function's return value. Just note that this works best if your function has no arguments (or arguments that are hashable):
from functools import cache import pandas as pd @cache def getdf2(): print("Running expensive DataFrame creation...") df = pd.DataFrame({"col1": range(10000), "col2": range(10000, 20000)}) return df.copy() # Return a copy to avoid external modifications breaking the cache
To clear the cache (if you need to refresh the DataFrame later), call getdf2.cache_clear().
3. Accidental Recomputation Triggers
Double-check if your function has any logic that forces a recompute every time—for example:
- Reading from a constantly updated file/database
- Using dynamic parameters that change on each call
- Modifying the DataFrame outside the function and expecting the cache to update
If any of these are happening, you'll need to adjust your caching logic to account for these changes (e.g., include the file's last modified time as a cache key, or add parameters to the function that trigger a recompute when changed).
内容的提问来源于stack exchange,提问作者user5944508

