如何判断Python导入语句对应的模块是内置还是用户自定义?
Got it, let's walk through a practical, step-by-step solution to tackle this problem—since you've got thousands of import statements in Excel, we'll need an automated approach using Python itself (makes sense, right?).
First, let's break down what we need to do: extract the imports from Excel, parse each line to get the root module, then categorize each module into built-in/standard library, user-defined, or third-party (since you might care about that too).
Step 1: Pull Import Statements from Excel
First, we'll use pandas to read the Excel file—super efficient for large datasets. If you don't have pandas installed, just run pip install pandas openpyxl (openpyxl handles .xlsx files).
import pandas as pd # Read your Excel file—adjust the column name to match your data df = pd.read_excel("your_imports_file.xlsx") import_lines = df["import_statement"].tolist() # Replace "import_statement" with your column header
Step 2: Parse Imports to Get Root Module Names
We need to extract the top-level module from each import line. For example:
from XYZ.loghelper import LogHelper→ root module isXYZimport os→ root module isosfrom .models import CustomUser→ this is a relative import (definitely user-defined)
Here's a regex-based helper function to do this:
import re def get_root_module(import_line): # Handle relative imports (start with .) if re.match(r'^from\s+\.', import_line): return "user_relative" # Marker for user-defined relative imports # Handle "import module" syntax import_match = re.match(r'^import\s+(\w+)', import_line) if import_match: return import_match.group(1) # Handle "from module.submodule import ..." syntax from_match = re.match(r'^from\s+(\w+)', import_line) if from_match: return from_match.group(1) # If we can't parse it, return None to mark as unknown return None
Step 3: Categorize Each Module
Now we need to figure out if a root module is:
- Built-in/Standard Library: Comes with Python (like
os,sys,json) - User-defined: Part of your project's codebase
- Third-party: Installed via pip (like
django,pandas)
First, Identify Standard Library Modules
We can precompute a list of all standard library modules to avoid importing each one (faster for large datasets):
import sys import os import importlib.util import pkgutil def get_standard_library_modules(): stdlib = set(sys.builtin_module_names) # Add built-in C extensions # Add pure-Python standard library modules python_install_path = os.path.dirname(os.path.dirname(sys.executable)) stdlib_lib_path = os.path.join(python_install_path, "lib") for _, module_name, _ in pkgutil.iter_modules(): try: spec = importlib.util.find_spec(module_name) if spec and spec.origin and spec.origin.startswith(stdlib_lib_path): stdlib.add(module_name) except: continue # Skip modules that cause errors return stdlib # Precompute once—this will take a few seconds but saves time later standard_lib_modules = get_standard_library_modules()
Next, Identify User-Defined Modules
You'll need to specify your project's root directory (where your code lives). We'll check if the module's file path is inside this directory.
PROJECT_ROOT = "/path/to/your/project" # Replace with your actual project path def is_user_defined(module_name): if module_name == "user_relative": return True # Relative imports are always user-defined try: spec = importlib.util.find_spec(module_name) if not spec or not spec.origin: return False # Check if the module's file is inside your project root return spec.origin.startswith(PROJECT_ROOT) except ModuleNotFoundError: # If the module isn't found, it might be a typo or missing—mark as unknown return False
Combine into a Classification Function
Now let's wrap this up into a function that takes an import line and returns its category:
def classify_import(import_line): root_module = get_root_module(import_line) if not root_module: return "Unknown" if root_module in standard_lib_modules: return "Built-in/Standard Library" if is_user_defined(root_module): return "User-defined" return "Third-party"
Step 4: Apply Classification to All Imports
Now we'll add the classification results back to the DataFrame and save it to a new Excel file:
# Handle lines with multiple imports (like "import os, sys") def split_multiple_imports(line): if "import" in line and "," in line: # Split into individual import statements parts = line.split("import")[1].split(",") return [f"import {part.strip()}" for part in parts] return [line] # Flatten the list of imports (split multi-import lines) all_imports = [] for line in import_lines: all_imports.extend(split_multiple_imports(line)) # Create a new DataFrame with cleaned imports clean_df = pd.DataFrame({"import_statement": all_imports}) # Add classification column clean_df["category"] = clean_df["import_statement"].apply(classify_import) # Save the result to a new Excel file clean_df.to_excel("classified_imports.xlsx", index=False)
Edge Cases to Watch For
- Relative Imports: We handled these explicitly since they're always part of your project.
- Aliased Imports: Like
import os as operating_system—the regex still capturesoscorrectly. - Typos/Missing Modules: These will show up as "Unknown"—you can manually review these rows.
- Third-Party Modules: If you don't care about separating these from user-defined, just adjust the classification function to merge them (but I included it since it's useful).
Testing the Workflow
Before running on all thousands of lines, test with a small sample of your imports to make sure the regex and classification logic works as expected. Adjust the regex or paths if needed for your specific cases.
内容的提问来源于stack exchange,提问作者prachi sutane

