You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python录屏同步麦克风录音及音视频合并技术方案问询

Hey there! Great job getting the screen capture and GUI display working with PIL. Let's walk through adding audio recording and merging it with your screen recording to make a full video with sound. Here's how to integrate everything step by step:

Step 1: Install Required Dependencies

First, you'll need a few extra libraries for audio capture and video processing. Run these commands in your terminal:

pip install pyaudio pillow opencv-python ffmpeg-python

Also, make sure FFmpeg is installed on your system and added to your PATH—you can download it from the official FFmpeg site, just ensure you can run ffmpeg from your terminal without errors.

Step 2: Core Components We’ll Add

  • Audio Capture: Use pyaudio to grab microphone input and save it as a WAV file.
  • Synchronized Control: Tie start/stop actions for both screen and audio recording to your GUI buttons.
  • Video Writing: Add code to save your captured screen frames to a video file (since you’re currently only displaying them).
  • Audio-Video Merging: Use FFmpeg to combine the screen video and recorded audio into a single polished output.

Step 3: Full Integrated Code

Assuming your existing code uses Tkinter for the GUI, here's the complete merged script with all features included:

import tkinter as tk
from PIL import Image, ImageTk, ImageGrab
import numpy as np
import cv2
import pyaudio
import wave
import threading
import time
import subprocess

# Global state variables for recording
is_recording = False
screen_thread = None
audio_thread = None
audio_stream = None
wave_file = None
video_writer = None

# Customize these settings for your setup
SCREEN_SIZE = (1920, 1080)  # Match your screen resolution
FPS = 20  # Adjust for smoothness vs file size

def start_recording():
    global is_recording, screen_thread, audio_thread, video_writer
    
    if is_recording:
        return
    
    is_recording = True
    
    # Initialize video writer to save screen frames
    fourcc = cv2.VideoWriter_fourcc(*'XVID')
    video_writer = cv2.VideoWriter('temp_screen.avi', fourcc, FPS, SCREEN_SIZE)
    
    # Start screen and audio recording in separate threads (to avoid freezing the GUI)
    screen_thread = threading.Thread(target=record_screen)
    screen_thread.start()
    
    audio_thread = threading.Thread(target=record_audio)
    audio_thread.start()
    
    # Update button states
    start_btn.config(text="Recording...", state=tk.DISABLED)
    stop_btn.config(state=tk.NORMAL)

def stop_recording():
    global is_recording, screen_thread, audio_thread, audio_stream, wave_file, video_writer
    
    if not is_recording:
        return
    
    is_recording = False
    
    # Wait for both threads to finish cleanly
    screen_thread.join()
    audio_thread.join()
    
    # Clean up resources
    if video_writer:
        video_writer.release()
    if audio_stream:
        audio_stream.stop_stream()
        audio_stream.close()
    if wave_file:
        wave_file.close()
    
    # Merge audio and video into the final output
    merge_audio_video()
    
    # Reset buttons
    start_btn.config(text="Start Recording", state=tk.NORMAL)
    stop_btn.config(state=tk.DISABLED)

def record_screen():
    global is_recording, video_writer
    
    while is_recording:
        # Capture screen with PIL
        screen_img = ImageGrab.grab(bbox=(0, 0, SCREEN_SIZE[0], SCREEN_SIZE[1]))
        # Convert to OpenCV's BGR format for video writing
        frame = cv2.cvtColor(np.array(screen_img), cv2.COLOR_RGB2BGR)
        video_writer.write(frame)
        
        # Update GUI preview
        tk_img = ImageTk.PhotoImage(image=screen_img)
        screen_label.config(image=tk_img)
        screen_label.image = tk_img
        
        # Maintain consistent FPS
        time.sleep(1/FPS)

def record_audio():
    global audio_stream, wave_file
    
    # Audio configuration (standard 44.1kHz stereo)
    CHUNK = 1024
    FORMAT = pyaudio.paInt16
    CHANNELS = 2
    RATE = 44100
    AUDIO_OUTPUT = "temp_audio.wav"
    
    audio = pyaudio.PyAudio()
    audio_stream = audio.open(
        format=FORMAT,
        channels=CHANNELS,
        rate=RATE,
        input=True,
        frames_per_buffer=CHUNK
    )
    
    wave_file = wave.open(AUDIO_OUTPUT, 'wb')
    wave_file.setnchannels(CHANNELS)
    wave_file.setsampwidth(audio.get_sample_size(FORMAT))
    wave_file.setframerate(RATE)
    
    while is_recording:
        data = audio_stream.read(CHUNK)
        wave_file.writeframes(data)
    
    audio.terminate()

def merge_audio_video():
    # Use FFmpeg via subprocess (more reliable for some setups than ffmpeg-python)
    try:
        subprocess.run([
            'ffmpeg', '-i', 'temp_screen.avi', '-i', 'temp_audio.wav',
            '-c:v', 'copy', '-c:a', 'aac', '-strict', 'experimental',
            'final_recording.mp4'
        ], check=True)
        print("Success! Final video saved as final_recording.mp4")
    except subprocess.CalledProcessError:
        print("Error merging audio and video—make sure FFmpeg is installed correctly.")

# Build the GUI
root = tk.Tk()
root.title("Screen Recorder with Audio")

# Screen preview label
screen_label = tk.Label(root)
screen_label.pack()

# Control buttons
control_frame = tk.Frame(root, pady=10)
control_frame.pack()

start_btn = tk.Button(
    control_frame, text="Start Recording", command=start_recording,
    padx=20, pady=10, font=('Arial', 12)
)
start_btn.grid(row=0, column=0, padx=5)

stop_btn = tk.Button(
    control_frame, text="Stop Recording", command=stop_recording,
    padx=20, pady=10, font=('Arial', 12), state=tk.DISABLED
)
stop_btn.grid(row=0, column=1, padx=5)

root.mainloop()

Step 4: Key Adjustments & Tips

  • Dynamic Screen Size: Replace SCREEN_SIZE with pyautogui.size() (install pyautogui first) to auto-detect your screen resolution.
  • Audio Sync: If you notice a slight delay between audio and video, tweak the FPS value or CHUNK size in the audio settings to match your system's performance.
  • Error Handling: Add try/except blocks around resource cleanup (like closing streams) to handle unexpected stops gracefully.
  • File Cleanup: Add code to delete the temporary temp_screen.avi and temp_audio.wav files after merging if you don't need them.

Step 5: Testing

  1. Run the script and click "Start Recording"—it will capture your screen and microphone simultaneously.
  2. Click "Stop Recording" to end the session; the script will automatically merge the audio and video into final_recording.mp4.
  3. Open the output file to verify audio and video are synchronized and capture everything correctly.

内容的提问来源于stack exchange,提问作者amit9867

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 22:37:31