You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

iOS 18语音合成音频保存异常:杂音、低音调问题求助

解决iOS 18上文本转音频杂音、音调异常问题

问题根源分析

  1. 格式参数冲突:原代码里pcmBuffer.format.settings同时包含kAudioFileMP3Type(MP3文件类型)和kAudioFormatLinearPCM(PCM编码格式),这两个参数互斥。iOS 18对格式校验更严格,直接复用会导致编码逻辑混乱,出现杂音。
  2. 强制格式转换失真:初始化AVAudioFile时强制指定commonFormat: .pcmFormatInt16,但iOS 18中AVSpeechSynthesizer输出的PCM格式大概率是浮点型(如.pcmFormatFloat32),强制转换会破坏音频数据结构,引发音调异常。
  3. 语音实例nil未处理:AVSpeechSynthesisVoice(identifier: voiceIdentifier)返回nil时,没有备用逻辑,可能自动使用了与输出格式不匹配的默认语音。

修复方案

1. 分离音频格式参数

明确目标输出格式(如MP3/AAC),单独定义编码设置,不要复用语音合成的PCM输出设置:

// 标准AAC编码设置(兼容性优于MP3,iOS全设备支持)
let audioOutputSettings: [String: Any] = [
    AVFormatIDKey: kAudioFormatMPEG4AAC,
    AVSampleRateKey: 44100,
    AVNumberOfChannelsKey: 2,
    AVEncoderBitRateKey: 128000,
    AVEncoderAudioQualityKey: AVAudioQuality.medium.rawValue
]

2. 适配系统输出格式初始化文件

初始化AVAudioFile时,匹配语音合成输出buffer的格式属性,避免强制转换:

self.audioFile = try AVAudioFile(
    forWriting: saveToURL,
    settings: audioOutputSettings,
    commonFormat: pcmBuffer.format.commonFormat,
    interleaved: pcmBuffer.format.isInterleaved
)

3. 处理语音实例nil的兜底逻辑

添加备用语音选择,确保合成用的语音实例有效:

guard let voice = AVSpeechSynthesisVoice(identifier: voiceIdentifier) ?? AVSpeechSynthesisVoice(language: "zh-CN") else {
    fatalError("无法获取有效语音实例")
}
utterance.voice = voice

4. 完整修改后的核心方法

func synthesize(utteranceString: String, usingVoiceIdentifier voiceIdentifier: String, utteranceRate: Float, pitchMultiplier: Float, andSaveAudioFileTo saveToURL: URL, completionHandler: ((_ pcmBuffer: AVAudioPCMBuffer, _ audioFile: AVAudioFile?)->Void)? = nil) {
    print("#", #function, "voiceIdentifier: \(voiceIdentifier)", terminator: "\n")
    
    let utterance = AVSpeechUtterance(string: utteranceString)
    utterance.rate = utteranceRate
    utterance.pitchMultiplier = pitchMultiplier
 
    // 语音实例兜底处理
    guard let voice = AVSpeechSynthesisVoice(identifier: voiceIdentifier) ?? AVSpeechSynthesisVoice(language: "zh-CN") else {
        fatalError("无法获取有效语音实例")
    }
    utterance.voice = voice
    
    // 定义输出音频格式设置
    let audioOutputSettings: [String: Any] = [
        AVFormatIDKey: kAudioFormatMPEG4AAC,
        AVSampleRateKey: 44100,
        AVNumberOfChannelsKey: 2,
        AVEncoderBitRateKey: 128000,
        AVEncoderAudioQualityKey: AVAudioQuality.medium.rawValue
    ]
    
    synthesizer.write(utterance) { (buffer: AVAudioBuffer) in
        print("in closure")
        guard let pcmBuffer = buffer as? AVAudioPCMBuffer else {
            fatalError("未知buffer类型: \(buffer)")
        }
        if pcmBuffer.frameLength == 0 {
            print("合成完成,buffer长度为0")
        } else {
            if self.audioFile == nil {
                print("初始化音频文件")
                do {
                    self.audioFile = try AVAudioFile(
                        forWriting: saveToURL,
                        settings: audioOutputSettings,
                        commonFormat: pcmBuffer.format.commonFormat,
                        interleaved: pcmBuffer.format.isInterleaved
                    )
                } catch {
                    print("初始化音频文件失败: \(error.localizedDescription)")
                    completionHandler?(pcmBuffer, nil)
                    return
                }
            }
            do {
                try self.audioFile!.write(from: pcmBuffer)
                completionHandler?(pcmBuffer, self.audioFile)
            } catch {
                print("写入音频文件失败: \(error.localizedDescription)")
                completionHandler?(pcmBuffer,nil)
            }
        }
    }
}

额外排查建议

  • 先输出PCM格式文件验证语音合成是否正常,若PCM无问题,再排查编码环节的设置。
  • iOS 18中部分新语音的输出格式可能有变化,可通过print(pcmBuffer.format)查看实际输出格式,针对性调整参数。
  • 若坚持使用MP3格式,将AVFormatIDKey改为kAudioFormatMP3,但注意部分iOS设备对MP3硬件编码支持有限,可能出现兼容性问题。

内容的提问来源于stack exchange,提问作者daniel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 14:25:17