如何检测getUserMedia音视频流是否含音频及用户是否发声?
Nice work getting the camera stream up and running! To add audio detection—both checking if an audio track exists and detecting when the user is speaking—here's how you can extend your code:
1. Check if an Audio Track Exists
First, verifying that your media stream includes an audio track is straightforward. The MediaStream object has a getAudioTracks() method that returns an array of all active audio tracks. If the array length is greater than 0, you know an audio track is available.
You can add this check right after you receive the mediaStream:
// Check for audio track presence const audioTracks = mediaStream.getAudioTracks(); if (audioTracks.length === 0) { console.log("No audio track available (user may have denied mic access or no mic exists)"); return; } console.log("Audio track is ready!");
This is useful even though you requested audio in your constraints—users can still deny mic permissions, or devices might not have a microphone.
2. Detect User Speech Activity
To detect when the user is actually speaking, you'll need to use the Web Audio API to analyze the audio stream's volume. Here's how to integrate this into your existing code:
var constraints = { audio: { echoCancellation: true }, video: { width: 640, height: 480 } }; navigator.mediaDevices .getUserMedia(constraints) .then(function (mediaStream) { videoInput = document.querySelector("video"); videoInput.srcObject = mediaStream; videoInput.onloadedmetadata = function (e) { videoInput.play(); }; // 1. Check for audio track presence const audioTracks = mediaStream.getAudioTracks(); if (audioTracks.length === 0) { console.log("No audio track available"); return; } console.log("Audio track is ready!"); // 2. Set up audio analysis for speech detection let audioContext; let analyser; let dataArray; // Initialize audio analysis (needs user interaction to bypass browser autoplay restrictions) function initAudioAnalysis() { if (audioContext) return; // Create audio context (with Safari fallback) audioContext = new (window.AudioContext || window.webkitAudioContext)(); const source = audioContext.createMediaStreamSource(mediaStream); // Create analyser to read audio data analyser = audioContext.createAnalyser(); analyser.fftSize = 256; // Balances precision and performance (must be a power of 2) const bufferLength = analyser.frequencyBinCount; dataArray = new Uint8Array(bufferLength); // Connect the audio source to the analyser source.connect(analyser); // No need to connect to destination—we're just analyzing, not playing back // Start checking for audio activity checkSpeechActivity(); } // Check if the user is speaking by measuring average volume function checkSpeechActivity() { // Get frequency data from the analyser analyser.getByteFrequencyData(dataArray); // Calculate average volume across all frequency bins let averageVolume = 0; for (let i = 0; i < dataArray.length; i++) { averageVolume += dataArray[i]; } averageVolume = averageVolume / dataArray.length; // Adjust this threshold based on your environment (lower = more sensitive) const speechThreshold = 20; const isSpeaking = averageVolume > speechThreshold; console.log(isSpeaking ? "User is speaking!" : "No audio input detected"); // Repeat the check on the next frame requestAnimationFrame(checkSpeechActivity); } // Trigger audio analysis on first user click (bypasses browser autoplay rules) document.addEventListener('click', initAudioAnalysis, { once: true }); }) .catch(function (err) { console.log(err.name + ": " + err.message); });
Key Notes:
- Browser Autoplay Restrictions: Browsers block audio context initialization unless it's triggered by a user interaction (like a click). That's why we listen for a click event to start the analysis.
- Threshold Adjustment: The
speechThresholdvalue may need tuning—higher values work better in noisy environments, while lower values are more sensitive in quiet spaces. - Performance: The
fftSizecontrols how detailed the audio analysis is. Smaller values (like 128) are faster, while larger values (like 512) give more precise volume readings. - Cleanup: If you stop needing the analysis later, remember to call
audioContext.close()and cancel therequestAnimationFrameto avoid memory leaks.
内容的提问来源于stack exchange,提问作者gaurang

