如何将Python函数拆解得到的字节码指令序列重新编译?
Great question! Taking the instruction sequence from dis.Bytecode and turning it back into a working function involves two core steps: building raw bytecode bytes from the instructions, then constructing a new code object (and eventually a callable function) using that bytecode. Here's a step-by-step breakdown with practical code examples:
Step 1: Get the Instruction Sequence
First, let's start with your original setup to capture the instructions:
import dis import types def speak(): print("moo") # Convert the function into a list of dis.Instruction objects instructions = list(dis.Bytecode(speak))
Step 2: Generate Raw Bytecode Bytes
Each Instruction has an opcode (a single-byte operation code) and an optional argument (a 2-byte value for operations that reference constants, names, or variables). We need to convert these into a bytes object that Python's interpreter can execute:
bytecode = b'' for instr in instructions: # Add the opcode as a single byte bytecode += bytes([instr.opcode]) # If the instruction has an argument, append it as 2 little-endian bytes if instr.arg is not None: bytecode += instr.arg.to_bytes(2, byteorder='little')
Step 3: Create a New Code Object
Python functions rely on underlying code objects. We'll use types.CodeType to build a new one, reusing most metadata from your original speak function (like constants, names, and stack size) to ensure compatibility:
original_code = speak.__code__ new_code = types.CodeType( original_code.co_argcount, # Number of positional arguments original_code.co_posonlyargcount, # Positional-only argument count original_code.co_kwonlyargcount, # Keyword-only argument count original_code.co_nlocals, # Number of local variables original_code.co_stacksize, # Maximum stack space needed original_code.co_flags, # Function flags (e.g., generator status) bytecode, # Our generated bytecode original_code.co_consts, # Constants used (None, 'moo') original_code.co_names, # Names referenced (print) original_code.co_varnames, # Local variable names original_code.co_filename, # Source filename "generated_speak", # Name for the new function original_code.co_firstlineno, # First line number (for debugging) original_code.co_lnotab, # Line number table original_code.co_freevars, # Free variables (empty here) original_code.co_cellvars # Cell variables (empty here) )
Step 4: Wrap the Code Object into a Callable Function
Finally, use types.FunctionType to turn the code object into a usable function, passing the global namespace so it can access print:
generated_speak = types.FunctionType(new_code, globals()) # Test the generated function generated_speak() # Outputs: moo
Key Notes
- If you modify the instruction sequence (e.g., change a
LOAD_CONSTargument index), you'll need to update theco_conststuple to match the new index. - The
co_stacksizemust accurately reflect the maximum number of values on the stack during execution—reusing the original value is safe unless you alter the instruction flow. - This approach works smoothly for simple functions, but more complex functions (with closures, default arguments, or nested scopes) will require handling additional metadata in the
CodeTypeconstructor.
内容的提问来源于stack exchange,提问作者Right leg

