You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何判断Python导入语句对应的模块是内置还是用户自定义?

How to Classify Python Imports as Built-in or User-Defined (from Excel)

Got it, let's walk through a practical, step-by-step solution to tackle this problem—since you've got thousands of import statements in Excel, we'll need an automated approach using Python itself (makes sense, right?).

First, let's break down what we need to do: extract the imports from Excel, parse each line to get the root module, then categorize each module into built-in/standard library, user-defined, or third-party (since you might care about that too).

Step 1: Pull Import Statements from Excel

First, we'll use pandas to read the Excel file—super efficient for large datasets. If you don't have pandas installed, just run pip install pandas openpyxl (openpyxl handles .xlsx files).

import pandas as pd

# Read your Excel file—adjust the column name to match your data
df = pd.read_excel("your_imports_file.xlsx")
import_lines = df["import_statement"].tolist()  # Replace "import_statement" with your column header

Step 2: Parse Imports to Get Root Module Names

We need to extract the top-level module from each import line. For example:

  • from XYZ.loghelper import LogHelper → root module is XYZ
  • import os → root module is os
  • from .models import CustomUser → this is a relative import (definitely user-defined)

Here's a regex-based helper function to do this:

import re

def get_root_module(import_line):
    # Handle relative imports (start with .)
    if re.match(r'^from\s+\.', import_line):
        return "user_relative"  # Marker for user-defined relative imports
    
    # Handle "import module" syntax
    import_match = re.match(r'^import\s+(\w+)', import_line)
    if import_match:
        return import_match.group(1)
    
    # Handle "from module.submodule import ..." syntax
    from_match = re.match(r'^from\s+(\w+)', import_line)
    if from_match:
        return from_match.group(1)
    
    # If we can't parse it, return None to mark as unknown
    return None

Step 3: Categorize Each Module

Now we need to figure out if a root module is:

  1. Built-in/Standard Library: Comes with Python (like os, sys, json)
  2. User-defined: Part of your project's codebase
  3. Third-party: Installed via pip (like django, pandas)

First, Identify Standard Library Modules

We can precompute a list of all standard library modules to avoid importing each one (faster for large datasets):

import sys
import os
import importlib.util
import pkgutil

def get_standard_library_modules():
    stdlib = set(sys.builtin_module_names)  # Add built-in C extensions
    
    # Add pure-Python standard library modules
    python_install_path = os.path.dirname(os.path.dirname(sys.executable))
    stdlib_lib_path = os.path.join(python_install_path, "lib")
    
    for _, module_name, _ in pkgutil.iter_modules():
        try:
            spec = importlib.util.find_spec(module_name)
            if spec and spec.origin and spec.origin.startswith(stdlib_lib_path):
                stdlib.add(module_name)
        except:
            continue  # Skip modules that cause errors
    
    return stdlib

# Precompute once—this will take a few seconds but saves time later
standard_lib_modules = get_standard_library_modules()

Next, Identify User-Defined Modules

You'll need to specify your project's root directory (where your code lives). We'll check if the module's file path is inside this directory.

PROJECT_ROOT = "/path/to/your/project"  # Replace with your actual project path

def is_user_defined(module_name):
    if module_name == "user_relative":
        return True  # Relative imports are always user-defined
    
    try:
        spec = importlib.util.find_spec(module_name)
        if not spec or not spec.origin:
            return False
        
        # Check if the module's file is inside your project root
        return spec.origin.startswith(PROJECT_ROOT)
    except ModuleNotFoundError:
        # If the module isn't found, it might be a typo or missing—mark as unknown
        return False

Combine into a Classification Function

Now let's wrap this up into a function that takes an import line and returns its category:

def classify_import(import_line):
    root_module = get_root_module(import_line)
    
    if not root_module:
        return "Unknown"
    
    if root_module in standard_lib_modules:
        return "Built-in/Standard Library"
    
    if is_user_defined(root_module):
        return "User-defined"
    
    return "Third-party"

Step 4: Apply Classification to All Imports

Now we'll add the classification results back to the DataFrame and save it to a new Excel file:

# Handle lines with multiple imports (like "import os, sys")
def split_multiple_imports(line):
    if "import" in line and "," in line:
        # Split into individual import statements
        parts = line.split("import")[1].split(",")
        return [f"import {part.strip()}" for part in parts]
    return [line]

# Flatten the list of imports (split multi-import lines)
all_imports = []
for line in import_lines:
    all_imports.extend(split_multiple_imports(line))

# Create a new DataFrame with cleaned imports
clean_df = pd.DataFrame({"import_statement": all_imports})

# Add classification column
clean_df["category"] = clean_df["import_statement"].apply(classify_import)

# Save the result to a new Excel file
clean_df.to_excel("classified_imports.xlsx", index=False)

Edge Cases to Watch For

  • Relative Imports: We handled these explicitly since they're always part of your project.
  • Aliased Imports: Like import os as operating_system—the regex still captures os correctly.
  • Typos/Missing Modules: These will show up as "Unknown"—you can manually review these rows.
  • Third-Party Modules: If you don't care about separating these from user-defined, just adjust the classification function to merge them (but I included it since it's useful).

Testing the Workflow

Before running on all thousands of lines, test with a small sample of your imports to make sure the regex and classification logic works as expected. Adjust the regex or paths if needed for your specific cases.

内容的提问来源于stack exchange,提问作者prachi sutane

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:09:08