如何查询远程Hadoop服务器中已安装的各类库?
Hey there! Since you're working on a remote Hadoop server with Spark pre-installed and can't install new libraries, here are practical methods to list all available packages (including Keras, TensorFlow, and others you might need for neural network training):
Python Libraries (Keras, TensorFlow, etc.)
These methods target the Python ecosystem, where most ML libraries like Keras/TensorFlow live:
Use
pip listfor a formatted package list
Run this command in your terminal to get a clean, sorted list of installed packages and their versions. If your server uses multiple Python versions, specify the one you'll use for training (e.g.,python3 -m pip listinstead ofpip list):pip list # Or for Python 3 specifically python3 -m pip listIf you only want to see packages installed in your user directory (not system-wide), add the
--userflag:pip list --userUse
pip freezefor requirements-style output
This command outputs packages in a format that can be saved to a requirements file, which is great for noting exact versions:pip freeze # Or for Python 3 python3 -m pip freezeQuery directly in a Python shell
If you prefer checking within a Python session (e.g., when testing imports), run these lines:# List all installed packages with versions import pkg_resources for dist in pkg_resources.working_set: print(f"{dist.project_name}=={dist.version}") # Alternatively, list all available modules (includes built-in ones) help('modules')
Spark-Related Libraries
Since your server has Spark installed, you might also want to check dependencies available to Spark:
Check Spark's Java/JAR dependencies
If you need to verify Spark's built-in or added JAR libraries (e.g., for data connectors), launch the PySpark shell and run:# In PySpark shell sc.listJars()This will return a list of all JAR files loaded by Spark.
Check Python libraries available to Spark
Spark uses the Python interpreter configured in its settings. To confirm which packages are available to PySpark, run:# Use Spark's dedicated Python interpreter to list packages /path/to/spark/bin/python -m pip listReplace
/path/to/sparkwith the actual Spark installation directory on your server (you can find this withecho $SPARK_HOME).
Bonus: Conda Environments (if applicable)
If your server uses Anaconda or Miniconda, you can list packages with:
conda list
This will include all packages in the active Conda environment.
内容的提问来源于stack exchange,提问作者Varshini

