You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Python文件中提取函数名、文档字符串及函数体?

The cleanest and most robust way to solve this without regex or manual text parsing is to use Python's built-in ast module (Abstract Syntax Trees). AST parsing properly understands Python's syntax structure, so it handles edge cases like multi-line docstrings, nested indentation, and string literals inside function bodies that regex would struggle with.

Here's a complete implementation:

import ast

def extract_function_info(file_path):
    with open(file_path, 'r', encoding='utf-8') as f:
        source_code = f.read()
    
    tree = ast.parse(source_code)
    function_info = {}
    function_count = 1
    
    for node in ast.walk(tree):
        # To target only top-level functions, replace ast.walk(tree) with tree.body
        if isinstance(node, ast.FunctionDef):
            # Extract function name
            func_name = node.name
            
            # Extract docstring (handles multi-line automatically)
            func_doc = ast.get_docstring(node) or ""
            
            # Extract function body, excluding the docstring if present
            body_nodes = node.body
            # Check if the first body node is a docstring expression
            if (body_nodes 
                and isinstance(body_nodes[0], ast.Expr) 
                and isinstance(body_nodes[0].value, (ast.Constant, ast.Str))):
                body_nodes = body_nodes[1:]
            
            # Get original source code for each body statement
            func_body = []
            for body_node in body_nodes:
                segment = ast.get_source_segment(source_code, body_node)
                if segment:
                    func_body.append(segment)
            
            # Join body lines to preserve formatting
            func_body_str = '\n'.join(func_body)
            
            # Add to result dict
            function_info[function_count] = {
                'name': func_name,
                'docstring': func_doc,
                'body': func_body_str
            }
            function_count += 1
    
    return function_info

How to use it:

Call the function with your target .py file path:

info = extract_function_info("your_file.py")

Testing with your example:

For the sample code you provided, this will return exactly the dictionary structure you requested. Docstrings are extracted correctly (including multi-line ones), and the function body preserves indentation and newline characters as they appear in the original file.

Key notes:

  • Nested Functions: ast.walk will pick up functions inside other functions or classes. If you only want top-level functions, replace ast.walk(tree) with tree.body in the loop.
  • No Docstrings: Functions without docstrings will have an empty string for the docstring key.
  • Python Versions: ast.Constant works for Python 3.8+. For older versions, use ast.Str instead.
  • Encoding: The code uses utf-8 encoding—adjust if your file uses a different encoding.

This approach is far more reliable than regex because it doesn't rely on pattern matching; it directly parses the code's structure to extract the exact information you need.

内容的提问来源于stack exchange,提问作者R32415

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 21:37:51