调用OpenAI语音转写API遇格式错误问题求助
问题:从Telegram下载的M4A音频提交OpenAI转写时提示格式不支持
从Telegram API获取的.m4a音频文件,下载过程无报错,但提交到OpenAI语音转写接口时,始终返回Invalid file format错误,尽管m4a属于官方列出的支持格式列表。原音频文件可正常播放,怀疑下载环节导致文件损坏,但未收到任何报错提示。尝试过调整multipart表单写法、修改文件名等方式,问题仍未解决。
依赖配置(Cargo.toml)
[dependencies] warp = "0.3" tokio = { version = "1", features = ["full"] } serde = { version = "1.0", features = ["derive"] } serde_json = "1.0" dotenv = "0.15" reqwest = { version = "0.12", features = ["json", "multipart", "stream"] } teloxide = "0.12" log = "0.4.22" env_logger = "0.11.3" log4rs = "1.2.0" lazy_static = "1.4" anyhow = "1.0" uuid = { version = "1.9.1", features = ["v4"] } #for unique file names tokio-util = "0.7.11" mime_guess = "2.0.3"
核心函数代码
transcribe_audio函数
async fn transcribe_audio(openai_key: &str, file_name: &str, mime_type: Option<&str>) -> Result<String, anyhow::Error> { log::info!("Audio: step 4: in transcribe_audio."); let client = Client::new(); log::info!("File name: {}", file_name); // Open file log::info!("Audio: step 4 initializing: opening file"); let file_handle = tokio::fs::File::open(file_name).await .context("Failed to open the file")?; // Create a stream from the file let bytes_stream = tokio_util::codec::FramedRead::new(file_handle, tokio_util::codec::BytesCodec::new()); // Create the multipart form let file_part = reqwest::multipart::Part::stream(reqwest::Body::wrap_stream(bytes_stream)) .file_name(file_name.to_string()) // Use the original file name .mime_str(mime_type.expect("couldn't give it a mime type"))?; // Use the provided MIME type, or a default one let form = reqwest::multipart::Form::new() .text("model", "whisper-1") .part("file", file_part); log::info!("Audio: step 4: successfully made file_part"); log::info!("beginning to send request to transcriptions"); let response = client .post("https://api.openai.com/v1/audio/transcriptions") .header("Authorization", format!("Bearer {}", openai_key)) .multipart(form) .send() .await .context("Failed to send the request to OpenAI")?; if !response.status().is_success() { let status = response.status(); let text = response.text().await.unwrap_or_else(|_| String::from("Failed to read response text")); anyhow::bail!("Received non-200 status code ({}) from OpenAI: {}", status, text); } let response_json: serde_json::Value = response.json().await .context("Failed to parse the response from OpenAI")?; let transcription = response_json["text"] .as_str() .ok_or_else(|| anyhow::anyhow!("Transcription not found in response"))? .to_string(); Ok(transcription) }
download_file函数
async fn download_file(url: &str, file_id: &str, mime_type: Option<&str>) -> Result<String, anyhow::Error> { log::info!("Audio: step 3: in download_file fn"); let client = Client::new(); log::info!("Audio: step 3 initializing"); // Send POST request to the URL to GET the file_path let response = client .get(url) .send() .await .with_context(|| format!("Failed to send GET request to URL: {}", url))?; // Ensure the request was successful if !response.status().is_success() { let error_message = format!("Received non-200 status code ({}) when trying to access URL: {}", response.status(), url); log::error!("{}", error_message); anyhow::bail!(error_message); } // Determine the file extension based on the MIME type let file_extension = match mime_type { Some("audio/flac") => "flac", Some("audio/m4a") => "m4a", Some("audio/mp3") => "mp3", Some("audio/mp4") => "mp4", Some("audio/mpeg") => "mpeg", Some("audio/mpga") => "mpga", Some("audio/oga") => "oga", Some("audio/webm") => "webm", Some("audio/wav") => "wav", Some("audio/ogg") => "ogg", // Add more MIME types and their corresponding file extensions as needed _ => "unknown", }; let filename = format!("{}.{}", Uuid::new_v4(), file_extension); log::info!("Audio: step 3: in download_file. filename is {filename}"); // Create the file let mut file = File::create(&filename) .await .with_context(|| format!("Failed to create file: {}", filename))?; // Extract the response content let content = response .bytes() .await .with_context(|| "Failed to read content from response".to_string())?; // Write content to the file file.write_all(&content) .await .with_context(|| format!("Failed to write content to file: {}", filename))?; log::info!("Audio: step 3 completed successfully"); Ok(filename) }
handle_audio_message函数
async fn handle_audio_message(bot_token: &str, chat_id: &u64, audio: &Audio, openai_key: &str) -> Result<(), anyhow::Error> { log::info!("Audio: step 2: In handle_audio_message fn"); // Get the file path from Telegram using the get_file function let file_path_on_telegram = get_file(bot_token, &audio.file_id).await?; log::info!("Audio: step 2: In handle_audio_message. got file path"); log::info!("Audio: step 2: In handle_audio_message. File_path is {file_path_on_telegram}"); // Download the audio file from Telegram let file_url = format!("https://api.telegram.org/file/bot{}/{}", bot_token, &file_path_on_telegram); log::info!("Audio: step 3: about to download audio file"); let file_name = download_file(&file_url, &audio.file_id, audio.mime_type.as_deref()).await?; // Call OpenAI API to transcribe audio log::info!("Audio: step 4: about to transcribe the audio message"); let transcription = transcribe_audio(openai_key, &file_name, audio.mime_type.as_deref()).await?; log::info!("audio message transcribed to: {}", transcription); let bot = Client::new(); bot.post(&format!("https://api.telegram.org/bot{}/sendMessage", bot_token)) .json(&serde_json::json!({ "chat_id": chat_id, "text": transcription, })) .send() .await?; Ok(()) }
get_file函数
async fn get_file(bot_token: &str, file_id: &str) -> Result<String> { log::info!("Audio: step 2 initializing. in get_file right now"); let client = Client::new(); let res: Value = client.post(&format!("https://api.telegram.org/bot{}/getFile", bot_token)) .form(&[("file_id", file_id)]) .send() .await? .json() .await?; let file_path = res["result"]["file_path"].as_str().unwrap().to_string(); Ok(file_path) }
错误日志
2024-07-04T06:01:31.482835598+00:00 - INFO - Audio: step 2: In handle_audio_message fn 2024-07-04T06:01:31.482843609+00:00 - INFO - Audio: step 2 initializing. in get_file right now 2024-07-04T06:01:31.894554705+00:00 - INFO - Audio: step 2: In handle_audio_message. got file path 2024-07-04T06:01:31.894688175+00:00 - INFO - Audio: step 2: In handle_audio_message. File_path is music/file_0.m4a 2024-07-04T06:01:31.894730325+00:00 - INFO - Audio: step 3: about to download audio file 2024-07-04T06:01:31.894759266+00:00 - INFO - Audio: step 3: in download_file fn 2024-07-04T06:01:31.929841708+00:00 - INFO - Audio: step 3 initializing 2024-07-04T06:01:32.425059459+00:00 - INFO - Audio: step 3: in download_file. filename is 0a76013b-bafe-4bd1-9a42-613c116a8904.m4a 2024-07-04T06:01:32.548868885+00:00 - INFO - Audio: step 3 completed successfully 2024-07-04T06:01:32.548991745+00:00 - INFO - Audio: step 4: about to transcribe the audio message 2024-07-04T06:01:32.549003405+00:00 - INFO - Audio: step 4: in transcribe_audio. 2024-07-04T06:01:32.583884887+00:00 - INFO - File name: 0a76013b-bafe-4bd1-9a42-613c116a8904.m4a 2024-07-04T06:01:32.583975949+00:00 - INFO - Audio: step 4 initializing: opening file 2024-07-04T06:01:32.584274150+00:00 - INFO - Audio: step 4: successfully made file_part 2024-07-04T06:01:32.584331379+00:00 - INFO - beginning to send request to transcriptions 2024-07-04T06:01:32.998999770+00:00 - ERROR - Error handling message: Received non-200 status code (400 Bad Request) from OpenAI: { "error": { "message": "Invalid file format. Supported formats: ['flac', 'm4a', 'mp3', 'mp4', 'mpeg', 'mpga', 'oga', 'ogg', 'wav', 'webm']", "type": "invalid_request_error", "param": null, "code": null } }
内容的提问来源于stack exchange,提问作者Jamie Laden
相关产品推荐
相关产品推荐

