Python脚本性能分析:导入及__init__.py耗时过高求助
Alright, let's tackle your Python import performance issues head-on—since you've already done the hard work of profiling with cProfile and KCacheGrind, we can focus on concrete fixes for that 90% import overhead and cyclic imports you've spotted.
First: Are All Those __init__.py Files Necessary?
Starting with Python 3.3, namespace packages don't require an __init__.py to be recognized as a valid package. If your __init__.py files are empty, or only exist to "mark" a directory as a package without any actual code, you can safely delete them. Each unnecessary __init__.py adds tiny overhead as Python executes it during import, and those add up when you have lots of packages.
If your __init__.py files do have code, we'll optimize that next.
Optimize __init__.py Code to Cut Import Time
Most import overhead from __init__.py comes from doing too much work during the import phase. Here's how to fix that:
- Remove costly runtime logic: Don't put database connections, file reads, complex calculations, or global variable initialization in
__init__.py. Move these tasks into functions or methods that run only when they're actually needed, not when the package is first imported. - Use lazy loading for exposed APIs: If your
__init__.pyimports submodules to expose their classes/functions (likefrom .submodule import MyClass), replace that with lazy loading using__getattr__. This way, the submodule is only imported when someone actually accesses the API:# In your package's __init__.py def __getattr__(name): if name == "MyClass": from .submodule import MyClass return MyClass elif name == "utility_func": from .utils import utility_func return utility_func raise AttributeError(f"module {__name__} has no attribute {name}") - Trim redundant imports: Audit each
__init__.pyand remove any imports that aren't meant to be part of the package's public API. If a submodule is only used internally by other parts of the package, there's no need to import it in__init__.py.
Fix Cyclic Imports to Eliminate Redundant Work
Cyclic imports (Module A imports B, B imports A) don't just cause runtime errors—they force Python to reprocess import logic multiple times, adding significant overhead. Try these fixes:
- Refactor shared code: Extract the logic that both modules depend on into a separate, independent utility module. For example, if A and B both use the same helper function, move it to
utils.py, then have A and B import fromutilsinstead of each other. - Delay imports to function scope: Instead of importing a dependent module at the top of your file, import it inside the function that uses it. This avoids triggering the import during module initialization:
# In Module A def process_data(): from .module_b import helper_function # Use helper_function here - Handle type hint cycles gracefully: If the cycle comes from type annotations, use Python's deferred type evaluation. Add
from __future__ import annotationsat the top of your modules, which lets you use unquoted type hints even if the referenced class isn't yet defined. Alternatively, use string-based type hints (e.g.,def func(obj: "MyOtherClass") -> None:).
Bonus: Dig Deeper Into Import Times
For a more granular look at which imports are taking the longest, run your script with:
python -X importtime your_script.py
This will print a detailed breakdown of each module's import time, sorted by duration. It's a quick way to pinpoint exactly which package or submodule is dragging down your startup.
内容的提问来源于stack exchange,提问作者jim jarnac

