Repository Wiki
IAHispano/Applio

模型、索引与音高特征配置

本页说明 Applio 推理链路中模型文件、检索索引以及 F0(基频/音高)特征配置如何从入口参数传递到语音转换管线,并重点解释音高估计、自动调音和建议音高的实际实现。

Purpose and Scope

本页覆盖单文件推理与批量推理所共享的配置边界:model_path/pth_path 指向目标声线模型,index_path 指向可选的特征检索索引,f0_method、pitch、f0_autotune 与 proposed_pitch 控制音高特征生成。页面也涵盖 index_rate、protect 等会影响索引融合和音高保护的参数。

模型下载、训练、模型信息查看以及界面控件本身不在本页展开;应在对应的模型管理或训练页面中说明。这里仅追踪这些参数进入转换 API 后的真实控制流。

Overview

core.py 的 run_infer_script 是一个参数适配层:它接收面向应用的参数,重命名为转换管线使用的关键字参数,然后调用 import_voice_converter() 返回的 infer_pipeline.convert_audio(**kwargs)。模型路径和索引路径在这一层保持为独立配置;音高相关参数则原样传入管线。

在 Pipeline.get_f0 中,f0_method 选择 CREPE、CREPE-Tiny、RMVPE、FCPE 或 Swift 之一。得到原始 F0 后,代码按优先级执行自动调音或建议音高调整,否则只应用显式 pitch;最后把 F0 转换为 mel 标度并量化到 1–255 的 255 个粗粒度桶,同时保留未量化的 f0bak 供后续转换使用。

Architecture

Loading diagram...

该结构反映了源码中的职责分层:入口函数不实现音高算法,而是构造 kwargs;get_f0 负责估计、调整和量化;voice_conversion 接收量化 F0、原始 F0、索引及其融合率。model_path 和 index_path 作为转换输入被同时传递,但源码片段没有把索引加载过程放在 core.py 中,因此索引文件的具体格式和加载细节应以转换器实现为准。

Source: core.py Source: pipeline.py

参数进入转换管线

单文件入口

run_infer_script 的签名把配置分成四组:

  1. 输入输出与资源:input_path、output_path、pth_path、index_path。
  2. 音高与索引融合:pitch、index_rate、volume_envelope、protect、f0_method。
  3. F0 后处理:f0_autotune、f0_autotune_strength、proposed_pitch、proposed_pitch_threshold。
  4. 音频处理和输出:切分、清理、格式、embedder 以及后处理效果参数。

函数没有在入口处验证模型或索引路径,也没有捕获 convert_audio 的异常;它将参数放入字典后直接调用转换器。因此路径不存在、格式不兼容或底层模型初始化失败时,错误处理责任在下游转换实现,而不是这个适配函数。

python
1kwargs = { 2 "audio_input_path": input_path, 3 "audio_output_path": output_path, 4 "model_path": pth_path, 5 "index_path": index_path, 6 "volume_envelope": volume_envelope, 7 "pitch": pitch, 8 "index_rate": index_rate, 9 "protect": protect, 10 "f0_method": f0_method, 11 "split_audio": split_audio, 12 "f0_autotune": f0_autotune, 13 "f0_autotune_strength": f0_autotune_strength, 14 "proposed_pitch": proposed_pitch, 15 "proposed_pitch_threshold": proposed_pitch_threshold, 16} 17infer_pipeline = import_voice_converter() 18infer_pipeline.convert_audio(**kwargs)

Source: core.py

批量入口

run_batch_infer_script 使用同一组模型、索引和 F0 参数,但把输入和输出映射为 audio_input_paths 与 audio_output_path。这意味着批量模式不是另一套音高算法,而是另一种输入集合/输出目录的适配形式;文档或调用方应确保批量转换器确实接受这些关键字名称。

python
1kwargs = { 2 "audio_input_paths": input_folder, 3 "audio_output_path": output_folder, 4 "model_path": pth_path, 5 "index_path": index_path, 6 "pitch": pitch, 7 "index_rate": index_rate, 8 "volume_envelope": volume_envelope, 9 "protect": protect, 10 "f0_method": f0_method, 11 "split_audio": split_audio, 12 "f0_autotune": f0_autotune, 13 "f0_autotune_strength": f0_autotune_strength, 14 "proposed_pitch": proposed_pitch, 15 "proposed_pitch_threshold": proposed_pitch_threshold, 16}

Source: core.py

Sources

(2 files)
(root)
rvc/infer