如何匹配音频与文本时间节点?C# SpeechRecognizer能否获取词起始时间?
Hey there! Let's tackle your problem step by step—you're trying to get word-level timestamps from audio files, and the native C# SpeechRecognizer isn't cutting it. Here's what you need to know:
1. Does the native SpeechRecognizer support word timestamps?
Short answer: No, it doesn't. The System.Speech.Recognition library is designed primarily for converting speech to text, but it doesn't expose APIs to retrieve start/end times for individual words. If you're sticking with this library out of the box, you won't get the timestamp data you need.
2. Alternative tools & libraries to get word timestamps
Luckily, there are several solid options that work well with C# (or can be easily integrated into a C# project):
- Microsoft Azure Speech SDK (Cloud-based, high accuracy)
This is the most straightforward upgrade if you don't mind using cloud services. Azure's Speech SDK has first-class C# support and explicitly returns word-level timestamps in its recognition results. Each recognized word includes an Offset (start time in 100-nanosecond units) and Duration, so you can calculate start/end times easily.
Here's a quick code snippet to get you started:
using Microsoft.CognitiveServices.Speech; using Microsoft.CognitiveServices.Speech.Audio; // Initialize speech config with your Azure subscription key and region var speechConfig = SpeechConfig.FromSubscription("your-subscription-key", "your-region"); using var audioConfig = AudioConfig.FromWavFileInput("path-to-your-audio-file.wav"); using var speechRecognizer = new SpeechRecognizer(speechConfig, audioConfig); var result = await speechRecognizer.RecognizeOnceAsync(); if (result.Reason == ResultReason.RecognizedSpeech) { foreach (var word in result.Words) { // Convert 100-nanosecond units to milliseconds for readability var startTimeMs = word.Offset / 10000; var endTimeMs = (word.Offset + word.Duration) / 10000; Console.WriteLine($"Word: {word.Text} | Start: {startTimeMs}ms | End: {endTimeMs}ms"); } }
- Whisper.NET (Offline, open-source, high accuracy)
If you need offline processing, Whisper.NET (a .NET wrapper for OpenAI's Whisper model) is a fantastic choice. It's lightweight, easy to set up via NuGet, and directly supports word timestamps. Plus, Whisper's recognition accuracy is top-tier for most languages.
Example code:
using Whisper.net; using Whisper.net.Ggml; // Download the Whisper model (use "base" for a balance of speed and accuracy) var modelName = "base.en"; await DownloadModel(modelName, "./models"); using var whisperFactory = WhisperFactory.FromPath("./models/ggml-base.en.bin"); using var processor = whisperFactory.CreateBuilder() .WithInputFile("path-to-your-audio-file.wav") .WithWordTimestamps() // Enable word-level timestamp output .Build(); await foreach (var result in processor.ProcessAsync()) { foreach (var segment in result.Segments) { foreach (var word in segment.Words) { Console.WriteLine($"Word: {word.Text} | Start: {word.Start:F2}s | End: {word.End:F2}s"); } } } // Helper method to download the Whisper model async Task DownloadModel(string modelName, string downloadPath) { using var modelStream = await WhisperGgmlDownloader.GetGgmlModelAsync(GgmlType.Base); using var fileWriter = File.Create(Path.Combine(downloadPath, $"ggml-{modelName}.bin")); await modelStream.CopyToAsync(fileWriter); }
- CMU Sphinx (Offline, open-source, legacy option)
CMU Sphinx is a classic open-source speech recognition engine with .NET bindings (via SphinxSharp). It supports word timestamps but requires more setup—you'll need to download language models, acoustic models, and dictionaries for your target language. It's a good pick if you need full offline control and don't mind a bit of configuration overhead.
3. Final Recommendation
If cloud services are an option, go with Azure Speech SDK—it's the most polished and reliable choice for C#. For offline use, Whisper.NET is the modern, high-accuracy alternative. The native SpeechRecognizer just isn't built for word-level timestamp retrieval, so switching to one of these tools is your best bet.
内容的提问来源于stack exchange,提问作者bublebboy

