You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ElevenLabs API生成音频与网站效果差异及优化咨询

问题与解决方案

问题描述

我开发了一款基于AI的医院导诊应用,通过Gemini API生成导诊文本,再调用ElevenLabs API将文本转换为语音。但测试发现,代码生成的output.mp3音质远劣于ElevenLabs官网生成的音频,想咨询该问题的原因(是否为参数设置问题),同时寻求支持土耳其语的语音模型及参数优化建议,相关Python代码如下:

import absl.flags
import absl.app
import absl.logging
import google.generativeai as genai
import requests
import os
import pygame

# Disable unnecessary logs
absl.flags.FLAGS.stderrthreshold = "FATAL"

# Configure your API key
genai.configure(api_key="GEMİNİ_APİ") # Gemini API

# Initialize Pygame Mixer
pygame.mixer.init()

class Response:
    def text(self, prompt, question):
        """Sends the prompt and question to the Gemini API and returns the response in text format."""
        self.prompt = prompt
        self.question = question

        # Combine prompt and question
        full_question = f"{self.prompt}\n{self.question}"

        # Send a request to the Gemini API
        model = genai.GenerativeModel("gemini-1.5-flash")
        self.response = model.generate_content(full_question)

    def text_response(self):
        # Display the response on the screen
        print(self.response.text)

    def voice_response(self):
        url = "https://api.elevenlabs.io/v1/text-to-speech/68gbrBPLYTEZzIIJ0apU"  # Voice model API
        querystring = {"optimize_streaming_latency":"2"}

        payload = {
            "text": self.response.text,
            "voice_settings": {
                "stability": 0.35,
                "similarity_boost": 0.85,
                "style": 0.55
            }
        }
        headers = {
            "xi-api-key": "ELEVENLABS_APİ_KEY",
            "Content-Type": "application/json"
        }
        response_voice = requests.request("POST", url, json=payload, headers=headers, params=querystring)

        if response_voice.status_code == 200:
            with open("output.mp3", "wb") as file:
                file.write(response_voice.content)
            print("Audio successfully created and saved to 'output.mp3'.")

            # Play the audio file
            pygame.mixer.music.load("output.mp3")
            pygame.mixer.music.play()

            # Wait until the audio playback is complete
            while pygame.mixer.music.get_busy():
                pygame.time.Clock().tick(10)

        else:
            print(f"Error: {response_voice.status_code} - {response_voice.text}")

# Main function
def main(argv):
    prompt = """You are a friendly, polite, and respectful male employee responsible for guiding patients to the correct department and floor in a hospital.
      You don't talk about things you don't know.
      I will give you information about the departments and floors in the hospital. You will answer the questions asked to you based on this information!
      If a patient has a problem, help them, approach them with good intentions, share your feelings with them, and give them moral support.

      Departments on the 1st floor: Anesthesiology and Reanimation, Appointment making, Brain and Neurosurgery, and Pediatric Surgery.
      Directions to the departments on the 1st floor:

      1. Anesthesiology and Reanimation: Go straight through door A1, it is the last door on the right.
      2. Appointment making: You will see it immediately to the right of the entrance.
      3. Brain and Neurosurgery: It is the 2nd door on the left from door A2.
      4. Neurosurgery: Go straight through door C1, it is the last right door on the 1st left.
      5. Pediatric Surgery: Go through C1 and it is the first door on the right.

      Based on this information, guide the people who come to you and always remember to not ask for anything more after your answer!
      Your answers should not be too short, at least 3 lines.

      """
    question = input("Your question: ")
    response = Response()
    response.text(prompt, question)
    response.voice_response()

# Main program
if __name__ == '__main__':
    absl.app.run(main)

音质差异原因分析

  • 流媒体延迟优化参数影响:代码中设置了optimize_streaming_latency=2,该参数会通过压缩音频、简化模型推理来降低延迟,直接导致音质下降。ElevenLabs官网默认不开启此优化,这是音质差异的核心原因。
  • 语音参数匹配度不足:官网生成时会自动匹配当前语音模型的最优参数组合,而你代码中的stability、similarity_boost、style参数设置可能未适配所选模型,进一步加剧音质失真。
  • 音频播放环节损耗:pygame.mixer的音频解码或播放过程可能存在轻微损耗,但影响远小于前两点。

支持土耳其语的ElevenLabs语音模型

  • 男性土耳其语模型:ErXwobaYiN019PkySvjV,发音清晰自然,适合正式导诊场景。
  • 女性土耳其语模型:XrExE9yKIg1WjnnlVkGX,语气温和友好,贴合医患沟通场景。

参数优化建议

  1. 关闭流媒体延迟优化:将optimize_streaming_latency设置为0(默认值)或直接移除该参数,优先保障音质。
  2. 适配土耳其语模型的参数调整:
    • stability: 0.6-0.7,平衡语音稳定性与自然度,避免语调生硬或过度波动。
    • similarity_boost: 0.9,提升生成语音与模型原声的相似度,减少失真。
    • style: 0.3-0.4,控制语音风格化程度,保持自然的服务语气。
  3. 指定高音质输出格式:在请求payload中添加output_format: "mp3_22050_32",使用更高采样率和比特率的编码格式。
  4. 启用多语言模型:添加model_id: "eleven_multilingual_v2",该模型对土耳其语的发音优化更到位。

修改后的核心代码示例

调整voice_response方法如下:

def voice_response(self):
    # 替换为土耳其语男性模型ID
    url = "https://api.elevenlabs.io/v1/text-to-speech/ErXwobaYiN019PkySvjV"
    # 关闭流媒体延迟优化
    querystring = {"optimize_streaming_latency":"0"}

    payload = {
        "text": self.response.text,
        "model_id": "eleven_multilingual_v2",  # 指定多语言模型
        "output_format": "mp3_22050_32",  # 高音质输出格式
        "voice_settings": {
            "stability": 0.65,
            "similarity_boost": 0.9,
            "style": 0.35
        }
    }
    headers = {
        "xi-api-key": "ELEVENLABS_APİ_KEY",
        "Content-Type": "application/json"
    }
    response_voice = requests.request("POST", url, json=payload, headers=headers, params=querystring)

    if response_voice.status_code == 200:
        with open("output.mp3", "wb") as file:
            file.write(response_voice.content)
        print("Audio successfully created and saved to 'output.mp3'.")

        pygame.mixer.music.load("output.mp3")
        pygame.mixer.music.play()
        while pygame.mixer.music.get_busy():
            pygame.time.Clock().tick(10)
    else:
        print(f"Error: {response_voice.status_code} - {response_voice.text}")

内容的提问来源于stack exchange,提问作者Ahmet Musab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 02:24:53