如何在Kinect v1场景下识别语音中的形容词?技术方案求助
Hey there, I’ve been through similar Kinect v1 voice recognition headaches before, so let’s walk through actionable solutions to get that adjective extraction working—even with a regular mic if needed.
If you want to stick with Kinect first, let’s fix the distance and sentence-recognition issues:
Tweak Confidence Thresholds
The default recognition confidence is often too strict for distant speech. Lower it to catch more valid inputs while filtering obvious garbage:// Assuming you're using Kinect's SpeechRecognitionEngine recognizer.RecognitionConfidence = RecognitionConfidence.Medium;Build a Flexible Grammar for Sentence-Based Detection
Instead of only listing single adjectives, create a grammar that allows adjectives to appear anywhere in a sentence. UseSrgsWildcardto wrap your adjective list, so the recognizer ignores filler words:var adjRule = new SrgsRule("AdjectiveDetection"); // Add all target adjectives here var adjList = new SrgsOneOf("happy", "blue", "big", "soft", "bright"); // Allow any words before/after the adjective adjRule.Elements.Add(new SrgsSequence(new SrgsWildcard(), adjList, new SrgsWildcard())); adjRule.Scope = SrgsRuleScope.Public; var grammarDoc = new SrgsDocument(adjRule); recognizer.LoadGrammar(new Grammar(grammarDoc));In the
SpeechRecognizedevent, extract the matched adjective by checking the semantic result or parsing the recognized text.Boost Audio Quality for Distant Speech
Kinect v1’s array mic has beamforming and noise suppression—use them:var audioSource = kinectSensor.AudioSource; audioSource.NoiseSuppression = true; audioSource.AutomaticGainControlEnabled = true; // If you're tracking the user's face, point the mic beam at them if (faceTrackingResult != null) { audioSource.BeamAngle = faceTrackingResult.FaceRotation.Z; }
Kinect v1’s voice stack is outdated—ditching it for a regular mic + modern APIs will solve most distance and sentence-recognition problems:
Offline Solution with System.Speech + NLP
First, fix that "Grammar referenced by grammar not found" error:- Go to Settings > Time & Language > Speech and install the full language pack (including dictation support) for your target language.
- Use
DictationGrammarto capture full speech, then use an NLP library like SharpNLP to extract adjectives:
using System.Speech.Recognition; using SharpNLP.Tools.Parser; var recognizer = new SpeechRecognitionEngine(new CultureInfo("en-US")); recognizer.LoadGrammar(new DictationGrammar()); recognizer.SpeechRecognized += (sender, e) => { var spokenText = e.Result.Text; // Parse text to get part-of-speech tags var parser = new EnglishTreebankParser(); var parsedTree = parser.Parse(spokenText); // Extract words tagged as adjectives (POS tags starting with "JJ") var adjectives = parsedTree.GetAllNodes() .Where(node => node.Label.Value.StartsWith("JJ")) .Select(node => node.Value); // Display these adjectives on screen }; recognizer.SetInputToDefaultAudioDevice(); recognizer.RecognizeAsync(RecognizeMode.Multiple);Note: Install SharpNLP via NuGet and download its pre-trained model files.
Online Solution with Cloud Speech APIs
For the best accuracy (even with distant speech), use cloud services like Azure Speech or Google Cloud Speech-to-Text. These include built-in or integrated NLP tools to extract adjectives:using Azure.AI.TextAnalytics; using Microsoft.CognitiveServices.Speech; // Initialize Speech Recognizer var speechConfig = SpeechConfig.FromSubscription("your-sub-key", "your-region"); var speechRecognizer = new SpeechRecognizer(speechConfig); // Initialize Text Analytics for POS tagging var textAnalyticsConfig = new TextAnalyticsClientOptions(); var textAnalyticsClient = new TextAnalyticsClient( new Uri("your-text-analytics-endpoint"), new AzureKeyCredential("your-ta-key")); speechRecognizer.Recognized += async (sender, e) => { if (e.Result.Reason == ResultReason.RecognizedSpeech) { var spokenText = e.Result.Text; // Analyze part-of-speech var posResult = await textAnalyticsClient.AnalyzeSentimentAsync(spokenText); // Extract adjectives from POS tags var adjectives = posResult.Value.Phrases .Where(phrase => phrase.PartOfSpeech.Tag == PartOfSpeechTag.Adjective) .Select(phrase => phrase.Text); // Display adjectives } }; await speechRecognizer.StartContinuousRecognitionAsync();This approach handles background noise and distant speech far better than Kinect’s native tools.
Combine Kinect’s face/skeleton tracking to locate the user, then use a regular array mic (or even your laptop’s mic) to focus on their voice. Adjust the mic’s input sensitivity or use third-party beamforming software to target the user’s position—this gives you the best of both worlds.
内容的提问来源于stack exchange,提问作者j-maas

