使用Azure Speech to Text和C#生成时间戳的代码错误解决
Let's break down and fix those two errors you're running into:
Error 1: Cannot assign to 'RequestWordLevelTimestamps' because it is a 'method group'
The issue here is that RequestWordLevelTimestamps is a method on SpeechConfig, not a property you can assign a value to. You tried to set it like config.RequestWordLevelTimestamps = true;, which won't work. Instead, just call the method to enable word-level timestamps:
// Replace your original assignment with this method call config.RequestWordLevelTimestamps();
Alternatively, you can achieve the same effect using the property setter syntax (either approach works):
config.SetProperty(PropertyId.SpeechServiceRequest_RequestWordLevelTimestamps, "true");
Error 2: The name 'result' does not exist in the current context
The result variable you defined only exists inside the Recognized event handler lambda. When you try to access it later in your Main method, the compiler has no idea what you're referring to. To get the full JSON response with timestamps, you need to access it inside the Recognized event logic, where the result is available.
Fix: Move JSON extraction to the Recognized event
Update your Recognized event handler to pull the JSON result when a valid speech recognition occurs:
recognizer.Recognized += (s, e) => { var result = e.Result; Console.WriteLine($"Reason: {result.Reason.ToString()}"); if (result.Reason == ResultReason.RecognizedSpeech) { Console.WriteLine($"Final result: Text: {result.Text}."); // Grab the full JSON response with word-level timestamps here var json = result.Properties.GetProperty(PropertyId.SpeechServiceResponse_JsonResult); Console.WriteLine("\nFull JSON with word timestamps:"); Console.WriteLine(json); } };
Full Corrected Code
Here's the complete code with both fixes applied, plus a small addition to handle error details in the canceled event:
using System; using System.Threading.Tasks; using Microsoft.CognitiveServices.Speech; using Microsoft.CognitiveServices.Speech.Audio; namespace NEST { internal class NewBaseType { static async Task Main(string[] args) { // Initialize speech config with your credentials var config = SpeechConfig.FromSubscription("subscriptionkey", "region"); // Enable detailed output and word-level timestamps config.OutputFormat = OutputFormat.Detailed; config.RequestWordLevelTimestamps(); // Fixed error 1 // Load your audio file using (var audioInput = AudioConfig.FromWavFileInput("C:/Users/MichaelSchwartz/source/repos/AI-102-Process-Speech-master/transcribe_speech_to_text/media/Zoom_audio.wav")) using (var recognizer = new SpeechRecognizer(config, audioInput)) { recognizer.Recognizing += (s, e) => { Console.WriteLine($"RECOGNIZING: Text={e.Result.Text}"); }; recognizer.Recognized += (s, e) => { var result = e.Result; Console.WriteLine($"Reason: {result.Reason.ToString()}"); if (result.Reason == ResultReason.RecognizedSpeech) { Console.WriteLine($"Final result: Text: {result.Text}."); // Fixed error 2: Access result within the event handler var json = result.Properties.GetProperty(PropertyId.SpeechServiceResponse_JsonResult); Console.WriteLine("\nFull JSON response with word-level timestamps:"); Console.WriteLine(json); } }; recognizer.Canceled += (s, e) => { Console.WriteLine($"\n Canceled. Reason: {e.Reason.ToString()}, CanceledReason: {e.Reason}"); // Add error details for easier debugging if (e.Reason == CancellationReason.Error) { Console.WriteLine($"Error details: {e.ErrorDetails}"); } }; recognizer.SessionStarted += (s, e) => { Console.WriteLine("\n Session started event."); }; recognizer.SessionStopped += (s, e) => { Console.WriteLine("\n Session stopped event."); }; // Start continuous recognition await recognizer.StartContinuousRecognitionAsync().ConfigureAwait(false); Console.WriteLine("Press Enter to stop"); Console.ReadLine(); // Stop recognition await recognizer.StopContinuousRecognitionAsync().ConfigureAwait(false); } } } }
Quick Notes
- Make sure you're using the latest version of the Azure Cognitive Services Speech SDK for C# to avoid API mismatches.
- The JSON output will include a
Wordsarray, where each entry has the word text, its start offset (in 100-nanosecond units), and duration (also in 100-nanosecond units). You can parse this to get precise word-level timings. - For long audio files, consider using Azure's Batch Transcription API, which also supports word-level timestamps and is optimized for longer content.
内容的提问来源于stack exchange,提问作者Michael Schwartz

