如何在Python模块中隐藏中间计算变量,仅暴露指定输出?
Great question! In Python, there are a few idiomatic ways to hide internal module variables and only expose the ones you want, similar to the code block scoping you're used to in Ruby, Scala, or Java. Let's break down the most effective approaches for your src/data.py module:
1. Use the __all__ Module Variable (Standard Convention)
Python has a built-in convention for defining public exports: the __all__ list. When you define this variable at the end of your module, it tells Python exactly which names should be imported when someone uses from data import *. It also signals to tools (and other developers) which variables are intended for public use.
Here's how to modify your code with __all__:
import numpy as np import pandas as pd sha256_perf = ( pd.read_csv('data/hashbench-output.txt', sep='\t', na_filter=False) .query('Algorithm == "SHA256"') ) mean_throughput = sha256_perf['Throughput (MiB/s)'].mean() variance = sha256_perf['Error (± MiB/s)'] ** 2 total_variance = variance.sum() row_count = sha256_perf.shape[0] variance_of_mean = total_variance / (row_count ** 2) error_of_mean = variance_of_mean ** 0.5 sha256_summary = pd.DataFrame(data=[[mean_throughput, error_of_mean]]) sha256_summary.columns = ['Mean Throughput (MiB/s)', 'Error (± MiB/s)'] # Define public exports __all__ = ['sha256_perf', 'sha256_summary']
Note: While __all__ controls import * behavior, running dir(data) will still show internal variables. But this is the standard way to communicate public API boundaries in Python.
2. Encapsulate Internal Logic in a Private Function (Cleanest Approach)
For a more thorough hiding of intermediate variables, wrap all your data processing logic inside a private function (prefixed with an underscore, _) and only expose the final variables at the module level. This keeps all intermediate values confined to the function's scope, so they won't appear in the module's namespace at all.
Modified code using this approach:
import numpy as np import pandas as pd def _load_sha256_metrics(): # All intermediate variables are scoped to this function sha256_perf = ( pd.read_csv('data/hashbench-output.txt', sep='\t', na_filter=False) .query('Algorithm == "SHA256"') ) mean_throughput = sha256_perf['Throughput (MiB/s)'].mean() variance = sha256_perf['Error (± MiB/s)'] ** 2 total_variance = variance.sum() row_count = sha256_perf.shape[0] variance_of_mean = total_variance / (row_count ** 2) error_of_mean = variance_of_mean ** 0.5 sha256_summary = pd.DataFrame(data=[[mean_throughput, error_of_mean]]) sha256_summary.columns = ['Mean Throughput (MiB/s)', 'Error (± MiB/s)'] return sha256_perf, sha256_summary # Expose the final variables at module level sha256_perf, sha256_summary = _load_sha256_metrics() # Optional: Add __all__ to enforce public API __all__ = ['sha256_perf', 'sha256_summary']
This is my recommended approach because it:
- Completely hides intermediate variables from the module namespace
- Makes your code more reusable (you can call
_load_sha256_metrics()again if needed) - Improves readability by grouping related logic together
3. Prefix Internal Variables with Underscores (Conventional Privacy)
If you prefer not to refactor into a function, you can prefix all intermediate variables with a single underscore (e.g., _mean_throughput). This is a Python convention that signals "this variable is internal, don't use it directly".
While these variables will still appear in dir(data), they won't be imported with from data import *, and other developers will recognize them as private.
Example snippet:
import numpy as np import pandas as pd sha256_perf = ( pd.read_csv('data/hashbench-output.txt', sep='\t', na_filter=False) .query('Algorithm == "SHA256"') ) _mean_throughput = sha256_perf['Throughput (MiB/s)'].mean() _variance = sha256_perf['Error (± MiB/s)'] ** 2 _total_variance = _variance.sum() _row_count = sha256_perf.shape[0] _variance_of_mean = _total_variance / (_row_count ** 2) _error_of_mean = _variance_of_mean ** 0.5 sha256_summary = pd.DataFrame(data=[[_mean_throughput, _error_of_mean]]) sha256_summary.columns = ['Mean Throughput (MiB/s)', 'Error (± MiB/s)'] __all__ = ['sha256_perf', 'sha256_summary']
Key Notes
Python doesn't have true "private" variables like some other languages, but these conventions effectively communicate intent and control what's exposed to module users. The function encapsulation method is the most robust for hiding internal state, while __all__ is the standard way to define your public API.
内容的提问来源于stack exchange,提问作者David Moles

