You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

移动端浏览器实现按压屏幕录音、抬手停止功能

Press-to-Record with iOS Safari Compatibility (No GUI Libraries)

Hey there! I’ve dealt with exactly this iOS Safari recording headache before—let’s walk through a solid, native API-based solution that fits your press-to-record, lift-to-send flow without any bulky GUI libraries.

Core Approach: Use the Native MediaRecorder API

Forget third-party libraries with forced UI—the standard Web MediaRecorder API is your best bet here. It’s supported on iOS Safari 14.3+, Android Chrome, and all modern desktop browsers, and gives you full control over recording start/stop and raw audio data.

Step 1: Handle Microphone Permission (Critical for iOS)

iOS Safari requires microphone permission to be triggered by a user-initiated event (like your press/touch start). Perfectly aligns with your flow—we’ll request permission on the first press:

let mediaStream;
let recorder;
let audioChunks = [];

// Bind to your press events (adjust for touch/mouse)
document.addEventListener('touchstart', startRecording, { passive: false });
document.addEventListener('mousedown', startRecording);

// Bind to your release events
document.addEventListener('touchend', stopAndSendRecording);
document.addEventListener('mouseup', stopAndSendRecording);
document.addEventListener('mouseleave', stopAndSendRecording); // Catch drag-off cases

async function startRecording(e) {
  // Prevent accidental multiple recording starts
  if (recorder?.state === 'recording') return;

  try {
    // Request microphone access (only triggers on user interaction)
    mediaStream = await navigator.mediaDevices.getUserMedia({ audio: true });
    
    // Pick iOS-compatible MIME type
    // 'audio/mp4' works across most iOS versions; 'audio/webm' for newer builds
    const mimeType = MediaRecorder.isTypeSupported('audio/mp4') ? 'audio/mp4' : 'audio/webm';
    recorder = new MediaRecorder(mediaStream, { mimeType });

    // Collect audio chunks as recording progresses
    recorder.ondataavailable = (event) => {
      if (event.data.size > 0) {
        audioChunks.push(event.data);
      }
    };

    recorder.start();
  } catch (err) {
    console.error('Recording failed:', err);
    // Add user feedback here (e.g., "Please enable microphone access in settings")
  }
}

Step 2: Stop Recording & Send Audio Data

On release, stop the recorder, bundle the chunks into a Blob, and send it to your speech-to-text service. You can send the Blob directly or convert it to base64 if your API requires it:

async function stopAndSendRecording() {
  if (!recorder || recorder.state !== 'recording') return;

  recorder.stop();
  mediaStream.getTracks().forEach(track => track.stop()); // Clean up microphone stream

  // Wait for final audio chunk to be captured
  await new Promise(resolve => recorder.onstop = resolve);

  // Create a single audio Blob from collected chunks
  const audioBlob = new Blob(audioChunks, { type: recorder.mimeType });
  audioChunks = []; // Reset for next recording

  // Option 1: Send Blob directly to your API (most efficient)
  try {
    const response = await fetch('/your-speech-to-text-endpoint', {
      method: 'POST',
      body: audioBlob,
      headers: {
        'Content-Type': recorder.mimeType
      }
    });
    const transcriptionResult = await response.json();
    console.log('Transcription:', transcriptionResult);
    // Handle the transcribed text in your app here
  } catch (err) {
    console.error('Failed to send recording:', err);
  }

  // Option 2: Convert to Base64 if your API requires text-formatted audio
  // const reader = new FileReader();
  // reader.onload = () => {
  //   const base64Audio = reader.result.split(',')[1];
  //   // Send base64Audio to your API
  // };
  // reader.readAsDataURL(audioBlob);
}

Step 3: iOS Safari Specific Fixes & Notes

  • MIME Type Check: Always use MediaRecorder.isTypeSupported() to pick a format iOS recognizes—audio/mp4 is the most reliable across iOS versions.
  • User Interaction Rule: Never attempt to start recording without a user tap/press—iOS blocks this for privacy reasons.
  • Touch Event Passive Mode: For touchstart, set { passive: false } if you need to disable scrolling during recording.
  • Stream Cleanup: Always stop media tracks after recording to prevent microphone access leaks.

Fallback for Older iOS Versions (<14.3)

If you need to support iOS 13 or earlier, MediaRecorder isn’t available. Use the (deprecated but still functional) ScriptProcessorNode to capture raw PCM data:

// Simplified PCM capture example for older iOS
function startPCMRecording() {
  const audioContext = new (window.AudioContext || window.webkitAudioContext)();
  const source = audioContext.createMediaStreamSource(mediaStream);
  const processor = audioContext.createScriptProcessor(4096, 1, 1);

  source.connect(processor);
  processor.connect(audioContext.destination);

  processor.onaudioprocess = (event) => {
    const inputBuffer = event.inputBuffer;
    const channelData = inputBuffer.getChannelData(0);
    // Collect raw PCM data here, then encode to WAV/MP4 manually for sending
  };
}

Note: You’ll need to handle PCM-to-WAV/MP4 encoding yourself for this fallback.

Key Takeaways

  • Stick to native APIs to avoid GUI bloat and compatibility headaches.
  • Align permission requests with user interactions to satisfy iOS’s privacy constraints.
  • Always clean up media resources to prevent memory leaks.

内容的提问来源于stack exchange,提问作者Max Kenney

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 19:42:27