模型、索引与音高特征配置
本页说明 Applio 推理链路中模型文件、检索索引以及 F0(基频/音高)特征配置如何从入口参数传递到语音转换管线,并重点解释音高估计、自动调音和建议音高的实际实现。
Purpose and Scope
本页覆盖单文件推理与批量推理所共享的配置边界:model_path/pth_path 指向目标声线模型,index_path 指向可选的特征检索索引,f0_method、pitch、f0_autotune 与 proposed_pitch 控制音高特征生成。页面也涵盖 index_rate、protect 等会影响索引融合和音高保护的参数。
模型下载、训练、模型信息查看以及界面控件本身不在本页展开;应在对应的模型管理或训练页面中说明。这里仅追踪这些参数进入转换 API 后的真实控制流。
Overview
core.py 的 run_infer_script 是一个参数适配层:它接收面向应用的参数,重命名为转换管线使用的关键字参数,然后调用 import_voice_converter() 返回的 infer_pipeline.convert_audio(**kwargs)。模型路径和索引路径在这一层保持为独立配置;音高相关参数则原样传入管线。
在 Pipeline.get_f0 中,f0_method 选择 CREPE、CREPE-Tiny、RMVPE、FCPE 或 Swift 之一。得到原始 F0 后,代码按优先级执行自动调音或建议音高调整,否则只应用显式 pitch;最后把 F0 转换为 mel 标度并量化到 1–255 的 255 个粗粒度桶,同时保留未量化的 f0bak 供后续转换使用。
Architecture
该结构反映了源码中的职责分层:入口函数不实现音高算法,而是构造 kwargs;get_f0 负责估计、调整和量化;voice_conversion 接收量化 F0、原始 F0、索引及其融合率。model_path 和 index_path 作为转换输入被同时传递,但源码片段没有把索引加载过程放在 core.py 中,因此索引文件的具体格式和加载细节应以转换器实现为准。
Source: core.py Source: pipeline.py
参数进入转换管线
单文件入口
run_infer_script 的签名把配置分成四组:
- 输入输出与资源:
input_path、output_path、pth_path、index_path。 - 音高与索引融合:
pitch、index_rate、volume_envelope、protect、f0_method。 - F0 后处理:
f0_autotune、f0_autotune_strength、proposed_pitch、proposed_pitch_threshold。 - 音频处理和输出:切分、清理、格式、embedder 以及后处理效果参数。
函数没有在入口处验证模型或索引路径,也没有捕获 convert_audio 的异常;它将参数放入字典后直接调用转换器。因此路径不存在、格式不兼容或底层模型初始化失败时,错误处理责任在下游转换实现,而不是这个适配函数。
1kwargs = {
2 "audio_input_path": input_path,
3 "audio_output_path": output_path,
4 "model_path": pth_path,
5 "index_path": index_path,
6 "volume_envelope": volume_envelope,
7 "pitch": pitch,
8 "index_rate": index_rate,
9 "protect": protect,
10 "f0_method": f0_method,
11 "split_audio": split_audio,
12 "f0_autotune": f0_autotune,
13 "f0_autotune_strength": f0_autotune_strength,
14 "proposed_pitch": proposed_pitch,
15 "proposed_pitch_threshold": proposed_pitch_threshold,
16}
17infer_pipeline = import_voice_converter()
18infer_pipeline.convert_audio(**kwargs)Source: core.py
批量入口
run_batch_infer_script 使用同一组模型、索引和 F0 参数,但把输入和输出映射为 audio_input_paths 与 audio_output_path。这意味着批量模式不是另一套音高算法,而是另一种输入集合/输出目录的适配形式;文档或调用方应确保批量转换器确实接受这些关键字名称。
1kwargs = {
2 "audio_input_paths": input_folder,
3 "audio_output_path": output_folder,
4 "model_path": pth_path,
5 "index_path": index_path,
6 "pitch": pitch,
7 "index_rate": index_rate,
8 "volume_envelope": volume_envelope,
9 "protect": protect,
10 "f0_method": f0_method,
11 "split_audio": split_audio,
12 "f0_autotune": f0_autotune,
13 "f0_autotune_strength": f0_autotune_strength,
14 "proposed_pitch": proposed_pitch,
15 "proposed_pitch_threshold": proposed_pitch_threshold,
16}Source: core.py