Python录屏同步麦克风录音及音视频合并技术方案问询
Hey there! Great job getting the screen capture and GUI display working with PIL. Let's walk through adding audio recording and merging it with your screen recording to make a full video with sound. Here's how to integrate everything step by step:
Step 1: Install Required Dependencies
First, you'll need a few extra libraries for audio capture and video processing. Run these commands in your terminal:
pip install pyaudio pillow opencv-python ffmpeg-python
Also, make sure FFmpeg is installed on your system and added to your PATH—you can download it from the official FFmpeg site, just ensure you can run ffmpeg from your terminal without errors.
Step 2: Core Components We’ll Add
- Audio Capture: Use
pyaudioto grab microphone input and save it as a WAV file. - Synchronized Control: Tie start/stop actions for both screen and audio recording to your GUI buttons.
- Video Writing: Add code to save your captured screen frames to a video file (since you’re currently only displaying them).
- Audio-Video Merging: Use FFmpeg to combine the screen video and recorded audio into a single polished output.
Step 3: Full Integrated Code
Assuming your existing code uses Tkinter for the GUI, here's the complete merged script with all features included:
import tkinter as tk from PIL import Image, ImageTk, ImageGrab import numpy as np import cv2 import pyaudio import wave import threading import time import subprocess # Global state variables for recording is_recording = False screen_thread = None audio_thread = None audio_stream = None wave_file = None video_writer = None # Customize these settings for your setup SCREEN_SIZE = (1920, 1080) # Match your screen resolution FPS = 20 # Adjust for smoothness vs file size def start_recording(): global is_recording, screen_thread, audio_thread, video_writer if is_recording: return is_recording = True # Initialize video writer to save screen frames fourcc = cv2.VideoWriter_fourcc(*'XVID') video_writer = cv2.VideoWriter('temp_screen.avi', fourcc, FPS, SCREEN_SIZE) # Start screen and audio recording in separate threads (to avoid freezing the GUI) screen_thread = threading.Thread(target=record_screen) screen_thread.start() audio_thread = threading.Thread(target=record_audio) audio_thread.start() # Update button states start_btn.config(text="Recording...", state=tk.DISABLED) stop_btn.config(state=tk.NORMAL) def stop_recording(): global is_recording, screen_thread, audio_thread, audio_stream, wave_file, video_writer if not is_recording: return is_recording = False # Wait for both threads to finish cleanly screen_thread.join() audio_thread.join() # Clean up resources if video_writer: video_writer.release() if audio_stream: audio_stream.stop_stream() audio_stream.close() if wave_file: wave_file.close() # Merge audio and video into the final output merge_audio_video() # Reset buttons start_btn.config(text="Start Recording", state=tk.NORMAL) stop_btn.config(state=tk.DISABLED) def record_screen(): global is_recording, video_writer while is_recording: # Capture screen with PIL screen_img = ImageGrab.grab(bbox=(0, 0, SCREEN_SIZE[0], SCREEN_SIZE[1])) # Convert to OpenCV's BGR format for video writing frame = cv2.cvtColor(np.array(screen_img), cv2.COLOR_RGB2BGR) video_writer.write(frame) # Update GUI preview tk_img = ImageTk.PhotoImage(image=screen_img) screen_label.config(image=tk_img) screen_label.image = tk_img # Maintain consistent FPS time.sleep(1/FPS) def record_audio(): global audio_stream, wave_file # Audio configuration (standard 44.1kHz stereo) CHUNK = 1024 FORMAT = pyaudio.paInt16 CHANNELS = 2 RATE = 44100 AUDIO_OUTPUT = "temp_audio.wav" audio = pyaudio.PyAudio() audio_stream = audio.open( format=FORMAT, channels=CHANNELS, rate=RATE, input=True, frames_per_buffer=CHUNK ) wave_file = wave.open(AUDIO_OUTPUT, 'wb') wave_file.setnchannels(CHANNELS) wave_file.setsampwidth(audio.get_sample_size(FORMAT)) wave_file.setframerate(RATE) while is_recording: data = audio_stream.read(CHUNK) wave_file.writeframes(data) audio.terminate() def merge_audio_video(): # Use FFmpeg via subprocess (more reliable for some setups than ffmpeg-python) try: subprocess.run([ 'ffmpeg', '-i', 'temp_screen.avi', '-i', 'temp_audio.wav', '-c:v', 'copy', '-c:a', 'aac', '-strict', 'experimental', 'final_recording.mp4' ], check=True) print("Success! Final video saved as final_recording.mp4") except subprocess.CalledProcessError: print("Error merging audio and video—make sure FFmpeg is installed correctly.") # Build the GUI root = tk.Tk() root.title("Screen Recorder with Audio") # Screen preview label screen_label = tk.Label(root) screen_label.pack() # Control buttons control_frame = tk.Frame(root, pady=10) control_frame.pack() start_btn = tk.Button( control_frame, text="Start Recording", command=start_recording, padx=20, pady=10, font=('Arial', 12) ) start_btn.grid(row=0, column=0, padx=5) stop_btn = tk.Button( control_frame, text="Stop Recording", command=stop_recording, padx=20, pady=10, font=('Arial', 12), state=tk.DISABLED ) stop_btn.grid(row=0, column=1, padx=5) root.mainloop()
Step 4: Key Adjustments & Tips
- Dynamic Screen Size: Replace
SCREEN_SIZEwithpyautogui.size()(installpyautoguifirst) to auto-detect your screen resolution. - Audio Sync: If you notice a slight delay between audio and video, tweak the
FPSvalue orCHUNKsize in the audio settings to match your system's performance. - Error Handling: Add
try/exceptblocks around resource cleanup (like closing streams) to handle unexpected stops gracefully. - File Cleanup: Add code to delete the temporary
temp_screen.aviandtemp_audio.wavfiles after merging if you don't need them.
Step 5: Testing
- Run the script and click "Start Recording"—it will capture your screen and microphone simultaneously.
- Click "Stop Recording" to end the session; the script will automatically merge the audio and video into
final_recording.mp4. - Open the output file to verify audio and video are synchronized and capture everything correctly.
内容的提问来源于stack exchange,提问作者amit9867

