论文指标计算与基准测试
本页介绍 piano_transcription 仓库中用于复现论文报告指标的评测子系统:以 pytorch/calculate_score_for_paper.py 为核心的两阶段评测流水线(离线概率推理 infer_prob + 指标计算 calculate_metrics),以及训练期分段评测器 pytorch/evaluate.py 中的 SegmentEvaluator。基准测试覆盖 MAESTRO V2.0.0 与 MAPS 数据集的 test 划分。
Purpose and Scope
本页覆盖以下内容:
- 论文级指标计算:音符级 precision / recall / F1(含力度与偏移)、踏板级指标、帧级指标的计算方式,全部基于
mir_eval与sklearn.metrics。 - 两阶段评测流水线:为什么先把模型输出概率落盘为
.pkl,再基于落盘结果反复调阈值算指标。 - 基准测试入口:
runme.sh中对 MAESTRO / MAPS test 划分的评测命令,以及与 Google Onsets and Frames 系统对比时的onsets_frames后处理分支。 - 训练期分段评测:
SegmentEvaluator的 AP(average precision)与 MAE(mean absolute error)指标,用于训练过程中的监控,与论文最终指标互补。
本页不覆盖的内容(留给兄弟页面):
- 推理引擎
PianoTranscription(pytorch/inference.py)内部的分段前向、模型加载与 MIDI 写盘逻辑——参见推理相关页面。 - 后处理器
RegressionPostProcessor/OnsetsFramesPostProcessor的完整实现(位于utils/utilities.py)——本页只说明它们在评测中的调用契约。 - 训练循环、损失函数(
pytorch/losses.py)与数据打包(utils/features.py)。
Overview
论文(High-resolution Piano Transcription with Pedals by Regressing Precise Onsets and Offsets Times)报告的指标体系由两类评测共同支撑:
- 论文最终指标(song-level,整曲评测):由
calculate_score_for_paper.py计算。对每首曲子:- 音符指标:用
mir_eval.transcription_velocity.precision_recall_f1_overlap(含力度)或mir_eval.transcription.precision_recall_f1_overlap(不含力度)比对参考与预测的[onset, offset]区间和音高,起音容差 0.05 s,偏移按offset_ratio=0.2与最小容差 0.05 s 匹配。 - 踏板指标:把踏板事件视为"单一音高"的区间(
pitches=np.ones(...)),用mir_eval.transcription以 0.2 s 起音容差比对;同时计算踏板帧级 F1。 - 帧级指标:对
frame_output以阈值二值化后,与frame_roll逐帧比对,取 sklearn 的 precision / recall / F1(正类索引[1])。
- 音符指标:用
- 训练期分段指标(segment-level):由
evaluate.py的SegmentEvaluator在 mini-batch 上计算frame_ap、onset_macro_ap、offset_ap,以及对回归头(reg_onset/reg_offset/velocity/reg_pedal_*/pedal_frame)计算带掩码的 MAE。这些指标衡量的是概率/回归质量,而后处理成事件后的指标才与论文表格对齐。
关键设计动机:评测与推理解耦。infer_prob 的 docstring 明确写道——先把概率写盘 "This will reduce duplicate computation for later evaluation"。因为在调 onset_threshold / offset_threshold / frame_threshold 三个阈值(默认 [0.3, 0.3, 0.3])时,GPU 前向是最贵的部分;落盘后,ScoreCalculator 可以纯 CPU 并行地反复扫阈值,无需重新跑模型。
Architecture
评测子系统的整体结构如下(节点名均对应真实代码实体):
要点说明:
- 阶段 1 与阶段 2 之间唯一的契约是
probs目录下的.pkl文件与hdf5s目录下的.h5文件名(get_filename(hdf5_path))一一对应;ScoreCalculator通过遍历 hdf5 目录、按split属性过滤来枚举待评曲目,再去probs_dir找同名.pkl。 TargetProcessor在评测侧被复用:infer_prob用它把 MIDI 事件重新解码为参考事件(note_events/pedal_events,extend_pedal=True),保证参考标注与训练时使用的目标编码完全一致,避免评测口径漂移。SegmentEvaluator与论文指标互不依赖:它只消费forward_dataloader返回的output_dict(含各类概率/回归输出与对应 roll),用于训练日志,不写盘、不做事件后处理。
Core Flow
阶段 1:infer_prob —— 概率落盘
infer_prob(args) 遍历 workspace/hdf5s/<dataset> 下所有 .h5,只保留 hf.attrs['split'] == split 的曲目,对每首曲子执行:读波形 → 重建参考标注 → GPU 转写 → 把"模型输出概率 + 参考标注"打包进一个 total_dict 写成 .pkl。
1# Load audio
2audio = int16_to_float32(hf['waveform'][:])
3midi_events = [e.decode() for e in hf['midi_event'][:]]
4midi_events_time = hf['midi_event_time'][:]
5
6# Ground truths processor
7target_processor = TargetProcessor(
8 segment_seconds=len(audio) / sample_rate,
9 frames_per_second=frames_per_second, begin_note=begin_note,
10 classes_num=classes_num)
11
12# Get ground truths
13(target_dict, note_events, pedal_events) = \
14 target_processor.process(start_time=0,
15 midi_events_time=midi_events_time,
16 midi_events=midi_events, extend_pedal=True)
17
18ref_on_off_pairs = np.array([[event['onset_time'], event['offset_time']] for event in note_events])
19ref_midi_notes = np.array([event['midi_note'] for event in note_events])
20ref_velocity = np.array([event['velocity'] for event in note_events])
21
22# Transcribe
23transcribed_dict = transcriptor.transcribe(audio, midi_path=None)
24output_dict = transcribed_dict['output_dict']
25
26# Pack probabilites to dump
27total_dict = {key: output_dict[key] for key in output_dict.keys()}
28total_dict['frame_roll'] = target_dict['frame_roll']
29total_dict['ref_on_off_pairs'] = ref_on_off_pairs
30total_dict['ref_midi_notes'] = ref_midi_notes
31total_dict['ref_velocity'] = ref_velocity
32
33if 'pedal_frame_output' in output_dict.keys():
34 total_dict['ref_pedal_on_off_pairs'] = \
35 np.array([[event['onset_time'], event['offset_time']] for event in pedal_events])
36 total_dict['pedal_frame_roll'] = target_dict['pedal_frame_roll']
37
38prob_path = os.path.join(probs_dir, '{}.pkl'.format(get_filename(hdf5_path)))
39create_folder(os.path.dirname(prob_path))
40pickle.dump(total_dict, open(prob_path, 'wb'))Source: calculate_score_for_paper.py
设计意图:
segment_seconds=len(audio) / sample_rate把"整曲"当作一个分段交给TargetProcessor,使参考标注覆盖全曲;模型侧由PianoTranscription内部按segment_samples切块前向(细节见推理页面)。- 参考事件(
ref_on_off_pairs/ref_midi_notes/ref_velocity/ref_pedal_on_off_pairs)与帧级标注(frame_roll/pedal_frame_roll)一起落盘,使阶段 2 完全不需要再碰 HDF5 的 MIDI 内容,只剩"文件名 ↔ split"的枚举依赖。 - 仅当模型输出含
pedal_frame_output(即Note_pedal联合模型)时才打包踏板参考,兼容纯音符模型。
阶段 2:ScoreCalculator —— 并行逐曲指标计算
metrics(params) 的并行与聚合逻辑:
1# Calculate metrics in parallel
2with ProcessPoolExecutor() as exector:
3 results = exector.map(self.calculate_score_per_song, list_args)
4
5stats_list = list(results)
6stats_dict = {}
7for key in stats_list[0].keys():
8 stats_dict[key] = [e[key] for e in stats_list if key in e.keys()]
9
10return stats_dictSource: calculate_score_for_paper.py
设计意图:
- 用
[e[key] for e in stats_list if key in e.keys()]这种过滤式聚合(而不是直接e[key])是刻意的:某些曲目可能没有踏板事件(len(ref_pedal_on_off_pairs) == 0时根本不写入pedal_*键),过滤可避免 KeyError。 __call__(params)只返回np.mean(stats_dict['f1']),把ScoreCalculator变成一个可被超参搜索器直接调用的目标函数(注意该键名为'f1',而逐曲结果中实际写的是note_f1,接入外部搜索器时需留意此不一致)。
逐曲计算 calculate_score_per_song
帧级指标的二值化与长度对齐:
1if self.evaluate_frame:
2 frame_threshold = frame_threshold
3 y_pred = (np.sign(total_dict['frame_output'] - frame_threshold) + 1) / 2
4 y_pred[np.where(y_pred==0.5)] = 0
5 y_true = total_dict['frame_roll']
6 y_pred = y_pred[0 : y_true.shape[0], :]
7 y_true = y_true[0 : y_pred.shape[0], :]
8
9 tmp = metrics.precision_recall_fscore_support(y_true.flatten(), y_pred.flatten())
10 return_dict['frame_precision'] = tmp[0][1]
11 return_dict['frame_recall'] = tmp[1][1]
12 return_dict['frame_f1'] = tmp[2][1]Source: calculate_score_for_paper.py
关键实现细节:
np.sign(x - t)结果为{-1, 0, 1},(+1)/2映射为{0, 0.5, 1};随后显式把0.5(即恰等于阈值的帧)清零,保证二值化无歧义。- 由于模型输出帧数与标注帧数可能因整曲长度取整差异相差一帧,双向截断
y_pred = y_pred[0 : y_true.shape[0]]、y_true = y_true[0 : y_pred.shape[0]]取最小公共长度。 precision_recall_fscore_support返回(precision, recall, fbeta, support)数组,索引[1]取"正类(有声/有 onset)"的指标;flatten()把 帧×音高 摊平成一次总体统计(micro 风格)。
后处理与事件级指标:
1# Post processor
2if self.post_processor_type == 'regression':
3 post_processor = RegressionPostProcessor(self.frames_per_second,
4 classes_num=self.classes_num, onset_threshold=onset_threshold,
5 offset_threshold=offset_threshold,
6 frame_threshold=frame_threshold,
7 pedal_offset_threshold=self.pedal_offset_threshold)
8
9elif self.post_processor_type == 'onsets_frames':
10 post_processor = OnsetsFramesPostProcessor(self.frames_per_second,
11 classes_num=self.classes_num)
12
13# Post process piano note outputs to piano note and pedal events information
14(est_on_off_note_vels, est_pedal_on_offs) = \
15 post_processor.output_dict_to_note_pedal_arrays(output_dict)
16"""est_on_off_note_vels: (events_num, 4), the four columns are: [onset_time, offset_time, piano_note, velocity],
17est_pedal_on_offs: (pedal_events_num, 2), the two columns are: [onset_time, offset_time]"""
18
19# # Detect piano notes from output_dict
20est_on_offs = est_on_off_note_vels[:, 0 : 2]
21est_midi_notes = est_on_off_note_vels[:, 2]
22est_vels = est_on_off_note_vels[:, 3] * self.velocity_scaleSource: calculate_score_for_paper.py
后处理器的选择由命令行 --post_processor_type 决定(默认 regression):高分辨率系统用 RegressionPostProcessor;onsets_frames 仅为与 Google Onsets and Frames 系统公平对比而保留。注意 OnsetsFramesPostProcessor 分支不接受阈值参数(Google 基线使用其原版固定逻辑),而 RegressionPostProcessor 的 onset/offset/frame 三个阈值正是 metrics(params) 传进来的可调超参。
音符指标(含力度)与踏板指标:
1if self.velocity:
2 (note_precision, note_recall, note_f1, _) = (
3 mir_eval.transcription_velocity.precision_recall_f1_overlap(
4 ref_intervals=ref_on_off_pairs,
5 ref_pitches=note_to_freq(ref_midi_notes),
6 ref_velocities=total_dict['ref_velocity'],
7 est_intervals=est_on_offs,
8 est_pitches=note_to_freq(est_midi_notes),
9 est_velocities=est_vels,
10 onset_tolerance=self.onset_tolerance,
11 offset_ratio=self.offset_ratio,
12 offset_min_tolerance=self.offset_min_tolerance))Source: calculate_score_for_paper.py
1if self.pedal:
2 ref_pedal_on_off_pairs = output_dict['ref_pedal_on_off_pairs']
3
4 # Calculate pedal metrics
5 if len(ref_pedal_on_off_pairs) > 0:
6 pedal_precision, pedal_recall, pedal_f1, _ = \
7 mir_eval.transcription.precision_recall_f1_overlap(
8 ref_intervals=ref_pedal_on_off_pairs,
9 ref_pitches=np.ones(ref_pedal_on_off_pairs.shape[0]),
10 est_intervals=est_pedal_on_offs,
11 est_pitches=np.ones(est_pedal_on_offs.shape[0]),
12 onset_tolerance=0.2,
13 offset_ratio=self.pedal_offset_ratio,
14 offset_min_tolerance=self.pedal_offset_min_tolerance)
15
16 return_dict['pedal_precision'] = pedal_precision
17 return_dict['pedal_recall'] = pedal_recall
18 return_dict['pedal_f1'] = pedal_f1Source: calculate_score_for_paper.py
设计意图:
- 力度维度:
mir_eval.transcription_velocity在区间/音高匹配的基础上额外要求力度匹配,est_vels乘以self.velocity_scale(来自config.velocity_scale)把归一化力度还原到 0–128 量纲,与ref_velocity(MIDI 原始力度)可比。 - 踏板建模技巧:
mir_eval.transcription只认 (interval, pitch) 结构,这里把踏板当成"所有事件音高均为 1 Hz"的单一类别,起音容差放宽到 0.2 s(远大于音符的 0.05 s),因为踏板动作的时间精度天然低于音符 onset。 self.velocity/self.pedal/self.evaluate_frame是构造函数里硬编码的开关(均为True),calculate_metrics的 docstring 也说明 "Users may adjust the hyper-parameters in ScoreCalculator to evaluate with or without offset, velocity and pedals" —— 复现论文消融实验需直接改这些成员变量。
训练期 SegmentEvaluator
1if 'frame_output' in output_dict.keys():
2 statistics['frame_ap'] = metrics.average_precision_score(
3 output_dict['frame_roll'].flatten(),
4 output_dict['frame_output'].flatten(), average='macro')
5
6if 'onset_output' in output_dict.keys():
7 statistics['onset_macro_ap'] = metrics.average_precision_score(
8 output_dict['onset_roll'].flatten(),
9 output_dict['onset_output'].flatten(), average='macro')
10
11if 'reg_onset_output' in output_dict.keys():
12 """Mask indictes only evaluate where either prediction or ground truth exists"""
13 mask = (np.sign(output_dict['reg_onset_output'] + output_dict['reg_onset_roll'] - 0.01) + 1) / 2
14 statistics['reg_onset_mae'] = mae(output_dict['reg_onset_output'],
15 output_dict['reg_onset_roll'], mask)
16
17if 'velocity_output' in output_dict.keys():
18 """Mask indictes only evaluate where onset exists"""
19 statistics['velocity_mae'] = mae(output_dict['velocity_output'],
20 output_dict['velocity_roll'] / 128, output_dict['onset_roll'])Source: evaluate.py
掩码 MAE 的实现与意图:
1def mae(target, output, mask):
2 if mask is None:
3 return np.mean(np.abs(target - output))
4 else:
5 target *= mask
6 output *= mask
7 return np.sum(np.abs(target - output)) / np.clip(np.sum(mask), 1e-8, np.inf)Source: evaluate.py
- 对
reg_onset/reg_offset回归头,掩码是pred + roll - 0.01 > 0,即只在"预测或真值至少一方非零"的位置计算 MAE。这是高分辨率回归的核心评测口径:回归头只在 onset/offset 附近有意义,若对全 0 区域也算 MAE,指标会被大量简单 0-0 匹配虚高。 - 对
velocity头,掩码直接用onset_roll,即只在真实 onset 帧上评力度误差;真值除以 128 归一化。 pedal系列回归头(reg_pedal_onset_mae/reg_pedal_offset_mae/pedal_frame_mae)用mask=None的全量 MAE。- 所有统计值最终
np.around(..., decimals=4)后返回,直接进入训练日志。
Data Model
.pkl(阶段 1 产物,阶段 2 输入)的 total_dict 结构:
| 键 | 形状 | 来源 | 说明 |
|---|---|---|---|
frame_output | (frames, classes_num) | 模型 | 帧级激活概率 |
reg_onset_output / reg_offset_output / velocity_output 等 | (frames, classes_num) | 模型 | 各回归/分类头输出,随 model_type 变化 |
pedal_frame_output 等 | (frames,) | 模型 | 仅 Note_pedal 联合模型存在 |
frame_roll | (frames, classes_num) | TargetProcessor | 帧级真值 |
pedal_frame_roll | (frames,) | TargetProcessor | 踏板帧级真值(仅联合模型) |
ref_on_off_pairs | (events_num, 2) | note_events | 参考音符 [onset_time, offset_time](秒) |
ref_midi_notes | (events_num,) | note_events | 参考 MIDI 音高 |
ref_velocity | (events_num,) | note_events | 参考力度(0–128) |
ref_pedal_on_off_pairs | (pedal_events_num, 2) | pedal_events | 参考踏板区间(仅联合模型) |
后处理输出契约(output_dict_to_note_pedal_arrays):
est_on_off_note_vels:(events_num, 4),四列为[onset_time, offset_time, piano_note, velocity]est_pedal_on_offs:(pedal_events_num, 2),两列为[onset_time, offset_time]
calculate_score_per_song 产出的逐曲 return_dict 键:
| 键组 | 条件 | 计算方式 |
|---|---|---|
frame_precision/recall/f1 | evaluate_frame=True | sklearn,micro(flatten 后取正类 [1]) |
note_precision/recall/f1 | velocity=True → mir_eval.transcription_velocity;否则 mir_eval.transcription | onset 容差 0.05 s,offset_ratio 0.2,offset_min_tolerance 0.05 s |
pedal_precision/recall/f1 | pedal=True 且 len(ref_pedal_on_off_pairs) > 0 | mir_eval.transcription,pitches 全 1,onset 容差 0.2 s |
pedal_frame_precision/recall/f1 | 同上 | pedal_frame_output 以固定 0.5 阈值二值化后 sklearn |
Configuration Options
ScoreCalculator 硬编码超参(构造函数内)
| 选项 | 类型 | 默认值 | 说明 |
|---|---|---|---|
velocity | bool | True | 是否使用含力度的 mir_eval.transcription_velocity |
pedal | bool | True | 是否计算踏板指标 |
evaluate_frame | bool | True | 是否计算帧级 F1 |
onset_tolerance | float | 0.05 | 音符起音匹配容差(秒) |
offset_ratio | float / None | 0.2 | 偏移匹配按区间长度的比例;None 表示不评偏移 |
offset_min_tolerance | float | 0.05 | 偏移匹配最小容差(秒) |
pedal_offset_threshold | float | 0.2 | 传给 RegressionPostProcessor 的踏板偏移阈值 |
pedal_offset_ratio | float / None | 0.2 | 踏板偏移匹配比例 |
pedal_offset_min_tolerance | float | 0.05 | 踏板偏移最小容差 |
frames_per_second | int | 来自 config.frames_per_second | 帧率,决定时间-帧换算 |
classes_num | int | 来自 config.classes_num | 音高类别数 |
velocity_scale | int | 来自 config.velocity_scale | 力度反归一化系数 |
post_processor_type | str | 'regression' | 'regression' 或 'onsets_frames' |
calculate_metrics 的阈值默认值
| 参数 | 类型 | 默认值 | 说明 |
|---|---|---|---|
thresholds | list[float] | [0.3, 0.3, 0.3] | [onset_threshold, offset_threshold, frame_threshold],作用于后处理与帧二值化 |
命令行接口
| 子命令 | 参数 | 说明 |
|---|---|---|
infer_prob | --workspace, --model_type, --augmentation, --checkpoint_path, --dataset {maestro,maps}, --split, --post_processor_type(默认 regression), --cuda | GPU 推理落盘 .pkl |
calculate_metrics | --workspace, --model_type, --augmentation, --dataset {maestro,maps}, --split, --post_processor_type(默认 regression) | 纯 CPU 指标计算与均值打印 |
目录约定
probs_dir 由四个维度拼成,calculate_metrics 必须使用与 infer_prob 完全一致的四个取值才能找到落盘文件:
1probs_dir = os.path.join(workspace, 'probs',
2 'model_type={}'.format(model_type),
3 'augmentation={}'.format(augmentation), 'dataset={}'.format(dataset),
4 'split={}'.format(split))Source: calculate_score_for_paper.py
Usage Examples
基准测试完整流程(摘自 runme.sh 评测段落)
1# ============ Evaluate (optional) ============
2# Inference probability for evaluation
3python3 pytorch/calculate_score_for_paper.py infer_prob --workspace=$WORKSPACE --model_type='Note_pedal' --checkpoint_path=$NOTE_PEDAL_CHECKPOINT_PATH --augmentation='none' --dataset='maestro' --split='test' --cuda
4
5# Calculate metrics
6python3 pytorch/calculate_score_for_paper.py calculate_metrics --workspace=$WORKSPACE --model_type='Note_pedal' --augmentation='aug' --dataset='maestro' --split='test'
7python3 pytorch/calculate_score_for_paper.py calculate_metrics --workspace=$WORKSPACE --model_type='Note_pedal' --augmentation='aug' --dataset='maps' --split='test'Source: runme.sh
值得注意的不对称性:infer_prob 用 --augmentation='none'(推理当然不做数据增强),而 calculate_metrics 用 --augmentation='aug'(这是 probs_dir 的目录名约定,对应训练时该 checkpoint 所属的增强设置)。README 中训练段落使用 --augmentation='aug' 训练,因此评测命令也沿用 'aug' 作为目录键;post_processor_type 两阶段均默认 regression,保持一致。
结果输出格式
1t1 = time.time()
2stats_dict = score_calculator.metrics(thresholds)
3print('Time: {:.3f}'.format(time.time() - t1))
4
5for key in stats_dict.keys():
6 print('{}: {:.4f}'.format(key, np.mean(stats_dict[key])))Source: calculate_score_for_paper.py
对每个指标键打印跨曲算术平均,即论文表格中的数值口径(per-song 平均,而非全局 micro 汇总)。
API Reference
ScoreCalculator(hdf5s_dir, probs_dir, split, post_processor_type='regression')
构造评测器:保存 split、帧率、类别数、力度缩放、各项容差开关;遍历 hdf5s_dir 缓存 self.hdf5_paths。
Parameters:
hdf5s_dir(str): hdf5 数据集根目录,用于枚举曲目与 split 过滤probs_dir(str):infer_prob落盘的.pkl目录split(str): 目标划分,如'test'post_processor_type(str):'regression'(默认)或'onsets_frames'
ScoreCalculator.__call__(params): float
把实例变成"阈值 → 平均 F1"的目标函数。
Parameters: params (list[float]): [onset_threshold, offset_threshold, frame_threshold]
Returns: np.mean(stats_dict['f1'])
ScoreCalculator.metrics(params): dict
Parameters: 同上
Returns: dict[str, list[float]],键为指标名,值为逐曲列表。
并行行为: ProcessPoolExecutor 默认 worker 数(CPU 核数),逐曲并行;calculate_score_per_song 必须可 pickle(模块级方法,满足要求)。
ScoreCalculator.calculate_score_per_song(args): dict
Parameters: args (list): [n, hdf5_path, params]
Returns: 单曲指标字典 {'frame_precision', 'frame_recall', 'frame_f1', 'note_precision', 'note_recall', 'note_f1', (可选) 'pedal_precision', 'pedal_recall', 'pedal_f1', 'pedal_frame_precision', 'pedal_frame_recall', 'pedal_frame_f1'}
Throws / 失败路径:
.pkl不存在(infer_prob未跑或目录参数不一致)→FileNotFoundError- 纯音符模型但
self.pedal=True→output_dict['ref_pedal_on_off_pairs']抛KeyError
SegmentEvaluator(model, batch_size).evaluate(dataloader): dict
Parameters: model (object), batch_size (int)
Returns: 统计字典,键随模型输出头存在性而变:frame_ap、onset_macro_ap、offset_ap、reg_onset_mae、reg_offset_mae、velocity_mae、reg_pedal_onset_mae、reg_pedal_offset_mae、pedal_frame_mae,均四舍五入到 4 位小数。
mae(target, output, mask=None): float
带掩码平均绝对误差;mask=None 时退化为全量 np.mean(np.abs(target - output)),分母用 np.clip(np.sum(mask), 1e-8, np.inf) 防零除。
Failure Modes, Edge Cases & Concurrency
.pkl与.h5失配:ScoreCalculator以 hdf5 文件名定位 pkl(get_filename)。若infer_prob与calculate_metrics的四元组(model_type / augmentation / dataset / split)任一不同,会得到FileNotFoundError;这是两阶段流水线最主要的运维坑点。- 空踏板参考:
len(ref_pedal_on_off_pairs) == 0时整段踏板指标(含pedal_frame_*)被跳过,不写入return_dict;聚合端用if key in e.keys()过滤,均值只对"有踏板的曲目"计算——即pedal_f1的口径是"踏板曲目子集平均"。 - 帧数不齐:预测与真值帧数可能差 1(整曲长度取整),双向截断取公共长度;不做任何插值。
- 阈值恰等:
np.sign产生的0.5被显式置 0(偏向保守/低召回),帧级与踏板帧级路径都做了同样处理。 - 踏板帧级阈值固定 0.5:与音符帧级(可调
frame_threshold)不同,pedal_frame_output二值化阈值硬编码 0.5,不参与阈值搜索。 - 并发模型:阶段 2 是进程级并行(
ProcessPoolExecutor),无共享状态,calculate_score_per_song只读自身 pickle,天然无锁。阶段 1 是单进程串行循环(GPU 推理),逐曲打印进度。 - debug 开关:
metrics()里debug = False时直接并行执行;置True则串行调用calculate_score_per_song并打印曲目路径,便于排查单曲问题(list_args[0:]保持全集)。 __call__键名不一致:ScoreCalculator.__call__返回np.mean(stats_dict['f1']),而metrics()产出的键是note_f1等;把它直接当黑盒优化目标会得到KeyError,需自行适配。
Performance & Operational Notes
- 为什么两阶段:一次
infer_prob(GPU)产出概率后,阈值扫描、消融(velocity/pedal/frame 开关)、跨后处理器对比(regressionvsonsets_frames)都能在 CPU 上重复进行,无需重跑模型——这是整条流水线最重要的性能设计。 - 逐曲并行的收益:
mir_eval的区间匹配在长曲(MAESTRO 单曲可达 10+ 分钟、上万音符)上是 CPU 密集操作,进程池按曲粒度并行可近似线性扩展。 - I/O 布局:每曲一个
.pkl,天然支持按曲并行读取与断点续跑;重命名/移动probs目录即失效。 - 内存:
infer_prob会把整曲概率 + 全部参考标注驻留内存后一次性 pickle;超长曲目内存占用主要来自(frames, 88)级别的各输出头。
Extension Points
- 新指标:在
calculate_score_per_song末尾向return_dict追加键即可自动进入并行、聚合与打印链路(聚合端按键过滤,天然兼容部分曲目缺键)。 - 阈值搜索:
ScoreCalculator.__call__(params)已具备"参数向量 → 标量目标"形态,可直接接入hyperopt/optuna之类的搜索器(注意修正其内部读取的键名)。 - 对齐其他基线系统:新增后处理器只需在
calculate_score_per_song的post_processor_type分支中注册,并保证实现output_dict_to_note_pedal_arrays(output_dict) -> (est_on_off_note_vels, est_pedal_on_offs)契约。 - 新数据集:CLI 已将
--dataset约束为{maestro, maps};只要utils/features.py打包出的.h5含split属性、waveform、midi_event、midi_event_time,评测链路即可复用(打包细节见数据准备相关页面)。
Related Links
- pytorch/calculate_score_for_paper.py —— 论文指标计算入口(
infer_prob/calculate_metrics/ScoreCalculator) - pytorch/evaluate.py —— 训练期
SegmentEvaluator与mae - runme.sh —— MAESTRO/MAPS 基准评测命令
- README.md —— MAESTRO V2.0.0 数据集与工作区目录结构说明
- 推理引擎、后处理器与数据打包的实现细节分别属于推理、后处理与数据准备等兄弟页面。