You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3.7实现电脑音频转文本并自动输入的可行性与方案咨询

Hey there! Let's break down how to build this project step by step—since you already know the keyboard module, we'll focus on the other pieces you need to put it all together.

整体 Implementation Plan

Your project has three core parts:

  1. Capture the audio your computer is playing (not microphone input)
  2. Convert that short audio clip (single word) to text
  3. Use keyboard to auto-type the text

Let's dive into each part with actionable details and code snippets.

1. Capturing System Audio

To grab audio that your computer is outputting (like the "testing" playback), you'll need to access your system's "loopback" audio device. Here's how to do it with sounddevice (a user-friendly audio library):

First, install the library:

pip install sounddevice numpy wave

Step 1: Find your loopback device

Run this code to list all available audio devices—look for something like "Stereo Mix" (Windows), "BlackHole" (macOS, you'll need to install it first), or "pulse" (Linux):

import sounddevice as sd
print(sd.query_devices())

Note the device ID of your loopback device (e.g., 2).

Step 2: Detect when to record a word

We'll listen for audio volume spikes to trigger recording. When the volume goes above a threshold, we start recording; when it stays below the threshold for 1 second, we stop (assuming that's the end of a single word).

Here's a basic volume-checking function:

def is_audio_active(audio_data, threshold=0.01):
    # Calculate root mean square (RMS) to measure volume
    rms = np.sqrt(np.mean(audio_data**2))
    return rms > threshold

2. Speech-to-Text (STT) for Single Words

For offline, reliable STT that works well with single words, OpenAI Whisper is perfect—it's free, easy to use, and accurate.

First, install Whisper and its dependencies:

pip install openai-whisper
# You'll also need ffmpeg installed on your system (search for how to install it for your OS)

Load a lightweight model (the base.en model is fast enough for single words):

model = whisper.load_model("base.en")

Then, write a function to convert a recorded WAV file to text:

def audio_to_text(audio_file_path):
    result = model.transcribe(audio_file_path, language="en")
    # Return the transcribed text, stripped of extra whitespace
    return result["text"].strip()

3. Auto-Type with keyboard

Since you already know this module, the part is straightforward—just use keyboard.write() to input the transcribed text. A small delay can help ensure the target window is active:

def auto_type_text(text):
    print(f"Auto-typing: {text}")
    time.sleep(0.5)  # Give yourself time to switch to the input window
    keyboard.write(text)

4. Putting It All Together

Here's a full code framework that ties everything together:

import sounddevice as sd
import numpy as np
import whisper
import keyboard
import wave
import time

# Configuration
SAMPLE_RATE = 16000  # Whisper's preferred sample rate
LOOPBACK_DEVICE_ID = 2  # Replace with your device ID from earlier
VOLUME_THRESHOLD = 0.01
SILENCE_TIMEOUT = 1.0  # Seconds of silence to end recording
TEMP_AUDIO_FILE = "temp_word.wav"

# Initialize components
model = whisper.load_model("base.en")
sd.default.samplerate = SAMPLE_RATE
sd.default.device = LOOPBACK_DEVICE_ID

def is_audio_active(audio_data):
    rms = np.sqrt(np.mean(audio_data**2))
    return rms > VOLUME_THRESHOLD

def save_audio_to_wav(audio_data, file_path):
    with wave.open(file_path, 'wb') as wf:
        wf.setnchannels(1)  # Mono audio
        wf.setsampwidth(2)  # 16-bit
        wf.setframerate(SAMPLE_RATE)
        wf.writeframes((audio_data * 32767).astype(np.int16))

def main():
    print("Listening for audio...")
    recording = False
    audio_buffer = []
    last_active_time = time.time()

    while True:
        # Capture a small chunk of audio
        chunk = sd.rec(int(SAMPLE_RATE * 0.1), channels=1, dtype=np.float32)
        sd.wait()  # Wait for chunk to be recorded
        chunk = chunk.flatten()

        if is_audio_active(chunk):
            if not recording:
                print("Started recording...")
                recording = True
            audio_buffer.extend(chunk)
            last_active_time = time.time()
        elif recording:
            # Check if we've been silent long enough
            if time.time() - last_active_time > SILENCE_TIMEOUT:
                print("Stopped recording.")
                recording = False
                # Save the recorded audio
                save_audio_to_wav(np.array(audio_buffer), TEMP_AUDIO_FILE)
                # Convert to text
                text = audio_to_text(TEMP_AUDIO_FILE)
                if text:
                    auto_type_text(text)
                # Reset buffer
                audio_buffer = []
        time.sleep(0.05)

if __name__ == "__main__":
    try:
        main()
    except KeyboardInterrupt:
        print("\nStopped listening.")

Key Notes to Tweak:

  • Adjust VOLUME_THRESHOLD based on your system—if you get false triggers, raise it; if it misses words, lower it.
  • For macOS, you'll need to install BlackHole (a virtual audio loopback device) to capture system audio.
  • If you want faster STT, use the tiny.en Whisper model instead of base.en (tradeoff is slightly lower accuracy).

内容的提问来源于stack exchange,提问作者Pepe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 14:02:56