Repository Wiki
bytedance/piano_transcription

论文指标计算与基准测试

本页介绍 piano_transcription 仓库中用于复现论文报告指标的评测子系统:以 pytorch/calculate_score_for_paper.py 为核心的两阶段评测流水线(离线概率推理 infer_prob + 指标计算 calculate_metrics),以及训练期分段评测器 pytorch/evaluate.py 中的 SegmentEvaluator。基准测试覆盖 MAESTRO V2.0.0 与 MAPS 数据集的 test 划分。

Purpose and Scope

本页覆盖以下内容:

  • 论文级指标计算:音符级 precision / recall / F1(含力度与偏移)、踏板级指标、帧级指标的计算方式,全部基于 mir_eval 与 sklearn.metrics。
  • 两阶段评测流水线:为什么先把模型输出概率落盘为 .pkl,再基于落盘结果反复调阈值算指标。
  • 基准测试入口:runme.sh 中对 MAESTRO / MAPS test 划分的评测命令,以及与 Google Onsets and Frames 系统对比时的 onsets_frames 后处理分支。
  • 训练期分段评测:SegmentEvaluator 的 AP(average precision)与 MAE(mean absolute error)指标,用于训练过程中的监控,与论文最终指标互补。

本页不覆盖的内容(留给兄弟页面):

  • 推理引擎 PianoTranscription(pytorch/inference.py)内部的分段前向、模型加载与 MIDI 写盘逻辑——参见推理相关页面。
  • 后处理器 RegressionPostProcessor / OnsetsFramesPostProcessor 的完整实现(位于 utils/utilities.py)——本页只说明它们在评测中的调用契约。
  • 训练循环、损失函数(pytorch/losses.py)与数据打包(utils/features.py)。

Overview

论文(High-resolution Piano Transcription with Pedals by Regressing Precise Onsets and Offsets Times)报告的指标体系由两类评测共同支撑:

  1. 论文最终指标(song-level,整曲评测):由 calculate_score_for_paper.py 计算。对每首曲子:
    • 音符指标:用 mir_eval.transcription_velocity.precision_recall_f1_overlap(含力度)或 mir_eval.transcription.precision_recall_f1_overlap(不含力度)比对参考与预测的 [onset, offset] 区间和音高,起音容差 0.05 s,偏移按 offset_ratio=0.2 与最小容差 0.05 s 匹配。
    • 踏板指标:把踏板事件视为"单一音高"的区间(pitches=np.ones(...)),用 mir_eval.transcription 以 0.2 s 起音容差比对;同时计算踏板帧级 F1。
    • 帧级指标:对 frame_output 以阈值二值化后,与 frame_roll 逐帧比对,取 sklearn 的 precision / recall / F1(正类索引 [1])。
  2. 训练期分段指标(segment-level):由 evaluate.py 的 SegmentEvaluator 在 mini-batch 上计算 frame_ap、onset_macro_ap、offset_ap,以及对回归头(reg_onset / reg_offset / velocity / reg_pedal_* / pedal_frame)计算带掩码的 MAE。这些指标衡量的是概率/回归质量,而后处理成事件后的指标才与论文表格对齐。

关键设计动机:评测与推理解耦。infer_prob 的 docstring 明确写道——先把概率写盘 "This will reduce duplicate computation for later evaluation"。因为在调 onset_threshold / offset_threshold / frame_threshold 三个阈值(默认 [0.3, 0.3, 0.3])时,GPU 前向是最贵的部分;落盘后,ScoreCalculator 可以纯 CPU 并行地反复扫阈值,无需重新跑模型。

Architecture

评测子系统的整体结构如下(节点名均对应真实代码实体):

Loading diagram...

要点说明:

  • 阶段 1 与阶段 2 之间唯一的契约是 probs 目录下的 .pkl 文件与 hdf5s 目录下的 .h5 文件名(get_filename(hdf5_path))一一对应;ScoreCalculator 通过遍历 hdf5 目录、按 split 属性过滤来枚举待评曲目,再去 probs_dir 找同名 .pkl。
  • TargetProcessor 在评测侧被复用:infer_prob 用它把 MIDI 事件重新解码为参考事件(note_events / pedal_events,extend_pedal=True),保证参考标注与训练时使用的目标编码完全一致,避免评测口径漂移。
  • SegmentEvaluator 与论文指标互不依赖:它只消费 forward_dataloader 返回的 output_dict(含各类概率/回归输出与对应 roll),用于训练日志,不写盘、不做事件后处理。

Core Flow

阶段 1:infer_prob —— 概率落盘

infer_prob(args) 遍历 workspace/hdf5s/<dataset> 下所有 .h5,只保留 hf.attrs['split'] == split 的曲目,对每首曲子执行:读波形 → 重建参考标注 → GPU 转写 → 把"模型输出概率 + 参考标注"打包进一个 total_dict 写成 .pkl。

python
1# Load audio 2audio = int16_to_float32(hf['waveform'][:]) 3midi_events = [e.decode() for e in hf['midi_event'][:]] 4midi_events_time = hf['midi_event_time'][:] 5 6# Ground truths processor 7target_processor = TargetProcessor( 8 segment_seconds=len(audio) / sample_rate, 9 frames_per_second=frames_per_second, begin_note=begin_note, 10 classes_num=classes_num) 11 12# Get ground truths 13(target_dict, note_events, pedal_events) = \ 14 target_processor.process(start_time=0, 15 midi_events_time=midi_events_time, 16 midi_events=midi_events, extend_pedal=True) 17 18ref_on_off_pairs = np.array([[event['onset_time'], event['offset_time']] for event in note_events]) 19ref_midi_notes = np.array([event['midi_note'] for event in note_events]) 20ref_velocity = np.array([event['velocity'] for event in note_events]) 21 22# Transcribe 23transcribed_dict = transcriptor.transcribe(audio, midi_path=None) 24output_dict = transcribed_dict['output_dict'] 25 26# Pack probabilites to dump 27total_dict = {key: output_dict[key] for key in output_dict.keys()} 28total_dict['frame_roll'] = target_dict['frame_roll'] 29total_dict['ref_on_off_pairs'] = ref_on_off_pairs 30total_dict['ref_midi_notes'] = ref_midi_notes 31total_dict['ref_velocity'] = ref_velocity 32 33if 'pedal_frame_output' in output_dict.keys(): 34 total_dict['ref_pedal_on_off_pairs'] = \ 35 np.array([[event['onset_time'], event['offset_time']] for event in pedal_events]) 36 total_dict['pedal_frame_roll'] = target_dict['pedal_frame_roll'] 37 38prob_path = os.path.join(probs_dir, '{}.pkl'.format(get_filename(hdf5_path))) 39create_folder(os.path.dirname(prob_path)) 40pickle.dump(total_dict, open(prob_path, 'wb'))

Source: calculate_score_for_paper.py

设计意图:

  • segment_seconds=len(audio) / sample_rate 把"整曲"当作一个分段交给 TargetProcessor,使参考标注覆盖全曲;模型侧由 PianoTranscription 内部按 segment_samples 切块前向(细节见推理页面)。
  • 参考事件(ref_on_off_pairs / ref_midi_notes / ref_velocity / ref_pedal_on_off_pairs)与帧级标注(frame_roll / pedal_frame_roll)一起落盘,使阶段 2 完全不需要再碰 HDF5 的 MIDI 内容,只剩"文件名 ↔ split"的枚举依赖。
  • 仅当模型输出含 pedal_frame_output(即 Note_pedal 联合模型)时才打包踏板参考,兼容纯音符模型。

阶段 2:ScoreCalculator —— 并行逐曲指标计算

Loading diagram...

metrics(params) 的并行与聚合逻辑:

python
1# Calculate metrics in parallel 2with ProcessPoolExecutor() as exector: 3 results = exector.map(self.calculate_score_per_song, list_args) 4 5stats_list = list(results) 6stats_dict = {} 7for key in stats_list[0].keys(): 8 stats_dict[key] = [e[key] for e in stats_list if key in e.keys()] 9 10return stats_dict

Source: calculate_score_for_paper.py

设计意图:

  • 用 [e[key] for e in stats_list if key in e.keys()] 这种过滤式聚合(而不是直接 e[key])是刻意的:某些曲目可能没有踏板事件(len(ref_pedal_on_off_pairs) == 0 时根本不写入 pedal_* 键),过滤可避免 KeyError。
  • __call__(params) 只返回 np.mean(stats_dict['f1']),把 ScoreCalculator 变成一个可被超参搜索器直接调用的目标函数(注意该键名为 'f1',而逐曲结果中实际写的是 note_f1,接入外部搜索器时需留意此不一致)。

逐曲计算 calculate_score_per_song

帧级指标的二值化与长度对齐:

python
1if self.evaluate_frame: 2 frame_threshold = frame_threshold 3 y_pred = (np.sign(total_dict['frame_output'] - frame_threshold) + 1) / 2 4 y_pred[np.where(y_pred==0.5)] = 0 5 y_true = total_dict['frame_roll'] 6 y_pred = y_pred[0 : y_true.shape[0], :] 7 y_true = y_true[0 : y_pred.shape[0], :] 8 9 tmp = metrics.precision_recall_fscore_support(y_true.flatten(), y_pred.flatten()) 10 return_dict['frame_precision'] = tmp[0][1] 11 return_dict['frame_recall'] = tmp[1][1] 12 return_dict['frame_f1'] = tmp[2][1]

Source: calculate_score_for_paper.py

关键实现细节:

  • np.sign(x - t) 结果为 {-1, 0, 1},(+1)/2 映射为 {0, 0.5, 1};随后显式把 0.5(即恰等于阈值的帧)清零,保证二值化无歧义。
  • 由于模型输出帧数与标注帧数可能因整曲长度取整差异相差一帧,双向截断 y_pred = y_pred[0 : y_true.shape[0]]、y_true = y_true[0 : y_pred.shape[0]] 取最小公共长度。
  • precision_recall_fscore_support 返回 (precision, recall, fbeta, support) 数组,索引 [1] 取"正类(有声/有 onset)"的指标;flatten() 把 帧×音高 摊平成一次总体统计(micro 风格)。

后处理与事件级指标:

python
1# Post processor 2if self.post_processor_type == 'regression': 3 post_processor = RegressionPostProcessor(self.frames_per_second, 4 classes_num=self.classes_num, onset_threshold=onset_threshold, 5 offset_threshold=offset_threshold, 6 frame_threshold=frame_threshold, 7 pedal_offset_threshold=self.pedal_offset_threshold) 8 9elif self.post_processor_type == 'onsets_frames': 10 post_processor = OnsetsFramesPostProcessor(self.frames_per_second, 11 classes_num=self.classes_num) 12 13# Post process piano note outputs to piano note and pedal events information 14(est_on_off_note_vels, est_pedal_on_offs) = \ 15 post_processor.output_dict_to_note_pedal_arrays(output_dict) 16"""est_on_off_note_vels: (events_num, 4), the four columns are: [onset_time, offset_time, piano_note, velocity], 17est_pedal_on_offs: (pedal_events_num, 2), the two columns are: [onset_time, offset_time]""" 18 19# # Detect piano notes from output_dict 20est_on_offs = est_on_off_note_vels[:, 0 : 2] 21est_midi_notes = est_on_off_note_vels[:, 2] 22est_vels = est_on_off_note_vels[:, 3] * self.velocity_scale

Source: calculate_score_for_paper.py

后处理器的选择由命令行 --post_processor_type 决定(默认 regression):高分辨率系统用 RegressionPostProcessor;onsets_frames 仅为与 Google Onsets and Frames 系统公平对比而保留。注意 OnsetsFramesPostProcessor 分支不接受阈值参数(Google 基线使用其原版固定逻辑),而 RegressionPostProcessor 的 onset/offset/frame 三个阈值正是 metrics(params) 传进来的可调超参。

音符指标(含力度)与踏板指标:

python
1if self.velocity: 2 (note_precision, note_recall, note_f1, _) = ( 3 mir_eval.transcription_velocity.precision_recall_f1_overlap( 4 ref_intervals=ref_on_off_pairs, 5 ref_pitches=note_to_freq(ref_midi_notes), 6 ref_velocities=total_dict['ref_velocity'], 7 est_intervals=est_on_offs, 8 est_pitches=note_to_freq(est_midi_notes), 9 est_velocities=est_vels, 10 onset_tolerance=self.onset_tolerance, 11 offset_ratio=self.offset_ratio, 12 offset_min_tolerance=self.offset_min_tolerance))

Source: calculate_score_for_paper.py

python
1if self.pedal: 2 ref_pedal_on_off_pairs = output_dict['ref_pedal_on_off_pairs'] 3 4 # Calculate pedal metrics 5 if len(ref_pedal_on_off_pairs) > 0: 6 pedal_precision, pedal_recall, pedal_f1, _ = \ 7 mir_eval.transcription.precision_recall_f1_overlap( 8 ref_intervals=ref_pedal_on_off_pairs, 9 ref_pitches=np.ones(ref_pedal_on_off_pairs.shape[0]), 10 est_intervals=est_pedal_on_offs, 11 est_pitches=np.ones(est_pedal_on_offs.shape[0]), 12 onset_tolerance=0.2, 13 offset_ratio=self.pedal_offset_ratio, 14 offset_min_tolerance=self.pedal_offset_min_tolerance) 15 16 return_dict['pedal_precision'] = pedal_precision 17 return_dict['pedal_recall'] = pedal_recall 18 return_dict['pedal_f1'] = pedal_f1

Source: calculate_score_for_paper.py

设计意图:

  • 力度维度:mir_eval.transcription_velocity 在区间/音高匹配的基础上额外要求力度匹配,est_vels 乘以 self.velocity_scale(来自 config.velocity_scale)把归一化力度还原到 0–128 量纲,与 ref_velocity(MIDI 原始力度)可比。
  • 踏板建模技巧:mir_eval.transcription 只认 (interval, pitch) 结构,这里把踏板当成"所有事件音高均为 1 Hz"的单一类别,起音容差放宽到 0.2 s(远大于音符的 0.05 s),因为踏板动作的时间精度天然低于音符 onset。
  • self.velocity / self.pedal / self.evaluate_frame 是构造函数里硬编码的开关(均为 True),calculate_metrics 的 docstring 也说明 "Users may adjust the hyper-parameters in ScoreCalculator to evaluate with or without offset, velocity and pedals" —— 复现论文消融实验需直接改这些成员变量。

训练期 SegmentEvaluator

python
1if 'frame_output' in output_dict.keys(): 2 statistics['frame_ap'] = metrics.average_precision_score( 3 output_dict['frame_roll'].flatten(), 4 output_dict['frame_output'].flatten(), average='macro') 5 6if 'onset_output' in output_dict.keys(): 7 statistics['onset_macro_ap'] = metrics.average_precision_score( 8 output_dict['onset_roll'].flatten(), 9 output_dict['onset_output'].flatten(), average='macro') 10 11if 'reg_onset_output' in output_dict.keys(): 12 """Mask indictes only evaluate where either prediction or ground truth exists""" 13 mask = (np.sign(output_dict['reg_onset_output'] + output_dict['reg_onset_roll'] - 0.01) + 1) / 2 14 statistics['reg_onset_mae'] = mae(output_dict['reg_onset_output'], 15 output_dict['reg_onset_roll'], mask) 16 17if 'velocity_output' in output_dict.keys(): 18 """Mask indictes only evaluate where onset exists""" 19 statistics['velocity_mae'] = mae(output_dict['velocity_output'], 20 output_dict['velocity_roll'] / 128, output_dict['onset_roll'])

Source: evaluate.py

掩码 MAE 的实现与意图:

python
1def mae(target, output, mask): 2 if mask is None: 3 return np.mean(np.abs(target - output)) 4 else: 5 target *= mask 6 output *= mask 7 return np.sum(np.abs(target - output)) / np.clip(np.sum(mask), 1e-8, np.inf)

Source: evaluate.py

  • 对 reg_onset / reg_offset 回归头,掩码是 pred + roll - 0.01 > 0,即只在"预测或真值至少一方非零"的位置计算 MAE。这是高分辨率回归的核心评测口径:回归头只在 onset/offset 附近有意义,若对全 0 区域也算 MAE,指标会被大量简单 0-0 匹配虚高。
  • 对 velocity 头,掩码直接用 onset_roll,即只在真实 onset 帧上评力度误差;真值除以 128 归一化。
  • pedal 系列回归头(reg_pedal_onset_mae / reg_pedal_offset_mae / pedal_frame_mae)用 mask=None 的全量 MAE。
  • 所有统计值最终 np.around(..., decimals=4) 后返回,直接进入训练日志。

Data Model

.pkl(阶段 1 产物,阶段 2 输入)的 total_dict 结构:

键形状来源说明
frame_output(frames, classes_num)模型帧级激活概率
reg_onset_output / reg_offset_output / velocity_output 等(frames, classes_num)模型各回归/分类头输出,随 model_type 变化
pedal_frame_output 等(frames,)模型仅 Note_pedal 联合模型存在
frame_roll(frames, classes_num)TargetProcessor帧级真值
pedal_frame_roll(frames,)TargetProcessor踏板帧级真值(仅联合模型)
ref_on_off_pairs(events_num, 2)note_events参考音符 [onset_time, offset_time](秒)
ref_midi_notes(events_num,)note_events参考 MIDI 音高
ref_velocity(events_num,)note_events参考力度(0–128)
ref_pedal_on_off_pairs(pedal_events_num, 2)pedal_events参考踏板区间(仅联合模型)

后处理输出契约(output_dict_to_note_pedal_arrays):

  • est_on_off_note_vels: (events_num, 4),四列为 [onset_time, offset_time, piano_note, velocity]
  • est_pedal_on_offs: (pedal_events_num, 2),两列为 [onset_time, offset_time]

calculate_score_per_song 产出的逐曲 return_dict 键:

键组条件计算方式
frame_precision/recall/f1evaluate_frame=Truesklearn,micro(flatten 后取正类 [1])
note_precision/recall/f1velocity=True → mir_eval.transcription_velocity;否则 mir_eval.transcriptiononset 容差 0.05 s,offset_ratio 0.2,offset_min_tolerance 0.05 s
pedal_precision/recall/f1pedal=True 且 len(ref_pedal_on_off_pairs) > 0mir_eval.transcription,pitches 全 1,onset 容差 0.2 s
pedal_frame_precision/recall/f1同上pedal_frame_output 以固定 0.5 阈值二值化后 sklearn

Configuration Options

ScoreCalculator 硬编码超参(构造函数内)

选项类型默认值说明
velocityboolTrue是否使用含力度的 mir_eval.transcription_velocity
pedalboolTrue是否计算踏板指标
evaluate_frameboolTrue是否计算帧级 F1
onset_tolerancefloat0.05音符起音匹配容差(秒)
offset_ratiofloat / None0.2偏移匹配按区间长度的比例;None 表示不评偏移
offset_min_tolerancefloat0.05偏移匹配最小容差(秒)
pedal_offset_thresholdfloat0.2传给 RegressionPostProcessor 的踏板偏移阈值
pedal_offset_ratiofloat / None0.2踏板偏移匹配比例
pedal_offset_min_tolerancefloat0.05踏板偏移最小容差
frames_per_secondint来自 config.frames_per_second帧率,决定时间-帧换算
classes_numint来自 config.classes_num音高类别数
velocity_scaleint来自 config.velocity_scale力度反归一化系数
post_processor_typestr'regression''regression' 或 'onsets_frames'

calculate_metrics 的阈值默认值

参数类型默认值说明
thresholdslist[float][0.3, 0.3, 0.3][onset_threshold, offset_threshold, frame_threshold],作用于后处理与帧二值化

命令行接口

子命令参数说明
infer_prob--workspace, --model_type, --augmentation, --checkpoint_path, --dataset {maestro,maps}, --split, --post_processor_type(默认 regression), --cudaGPU 推理落盘 .pkl
calculate_metrics--workspace, --model_type, --augmentation, --dataset {maestro,maps}, --split, --post_processor_type(默认 regression)纯 CPU 指标计算与均值打印

目录约定

probs_dir 由四个维度拼成,calculate_metrics 必须使用与 infer_prob 完全一致的四个取值才能找到落盘文件:

python
1probs_dir = os.path.join(workspace, 'probs', 2 'model_type={}'.format(model_type), 3 'augmentation={}'.format(augmentation), 'dataset={}'.format(dataset), 4 'split={}'.format(split))

Source: calculate_score_for_paper.py

Usage Examples

基准测试完整流程(摘自 runme.sh 评测段落)

bash
1# ============ Evaluate (optional) ============ 2# Inference probability for evaluation 3python3 pytorch/calculate_score_for_paper.py infer_prob --workspace=$WORKSPACE --model_type='Note_pedal' --checkpoint_path=$NOTE_PEDAL_CHECKPOINT_PATH --augmentation='none' --dataset='maestro' --split='test' --cuda 4 5# Calculate metrics 6python3 pytorch/calculate_score_for_paper.py calculate_metrics --workspace=$WORKSPACE --model_type='Note_pedal' --augmentation='aug' --dataset='maestro' --split='test' 7python3 pytorch/calculate_score_for_paper.py calculate_metrics --workspace=$WORKSPACE --model_type='Note_pedal' --augmentation='aug' --dataset='maps' --split='test'

Source: runme.sh

值得注意的不对称性:infer_prob 用 --augmentation='none'(推理当然不做数据增强),而 calculate_metrics 用 --augmentation='aug'(这是 probs_dir 的目录名约定,对应训练时该 checkpoint 所属的增强设置)。README 中训练段落使用 --augmentation='aug' 训练,因此评测命令也沿用 'aug' 作为目录键;post_processor_type 两阶段均默认 regression,保持一致。

结果输出格式

python
1t1 = time.time() 2stats_dict = score_calculator.metrics(thresholds) 3print('Time: {:.3f}'.format(time.time() - t1)) 4 5for key in stats_dict.keys(): 6 print('{}: {:.4f}'.format(key, np.mean(stats_dict[key])))

Source: calculate_score_for_paper.py

对每个指标键打印跨曲算术平均,即论文表格中的数值口径(per-song 平均,而非全局 micro 汇总)。

API Reference

ScoreCalculator(hdf5s_dir, probs_dir, split, post_processor_type='regression')

构造评测器:保存 split、帧率、类别数、力度缩放、各项容差开关;遍历 hdf5s_dir 缓存 self.hdf5_paths。

Parameters:

  • hdf5s_dir (str): hdf5 数据集根目录,用于枚举曲目与 split 过滤
  • probs_dir (str): infer_prob 落盘的 .pkl 目录
  • split (str): 目标划分,如 'test'
  • post_processor_type (str): 'regression'(默认)或 'onsets_frames'

ScoreCalculator.__call__(params): float

把实例变成"阈值 → 平均 F1"的目标函数。

Parameters: params (list[float]): [onset_threshold, offset_threshold, frame_threshold]

Returns: np.mean(stats_dict['f1'])

ScoreCalculator.metrics(params): dict

Parameters: 同上

Returns: dict[str, list[float]],键为指标名,值为逐曲列表。

并行行为: ProcessPoolExecutor 默认 worker 数(CPU 核数),逐曲并行;calculate_score_per_song 必须可 pickle(模块级方法,满足要求)。

ScoreCalculator.calculate_score_per_song(args): dict

Parameters: args (list): [n, hdf5_path, params]

Returns: 单曲指标字典 {'frame_precision', 'frame_recall', 'frame_f1', 'note_precision', 'note_recall', 'note_f1', (可选) 'pedal_precision', 'pedal_recall', 'pedal_f1', 'pedal_frame_precision', 'pedal_frame_recall', 'pedal_frame_f1'}

Throws / 失败路径:

  • .pkl 不存在(infer_prob 未跑或目录参数不一致)→ FileNotFoundError
  • 纯音符模型但 self.pedal=True → output_dict['ref_pedal_on_off_pairs'] 抛 KeyError

SegmentEvaluator(model, batch_size).evaluate(dataloader): dict

Parameters: model (object), batch_size (int)

Returns: 统计字典,键随模型输出头存在性而变:frame_ap、onset_macro_ap、offset_ap、reg_onset_mae、reg_offset_mae、velocity_mae、reg_pedal_onset_mae、reg_pedal_offset_mae、pedal_frame_mae,均四舍五入到 4 位小数。

mae(target, output, mask=None): float

带掩码平均绝对误差;mask=None 时退化为全量 np.mean(np.abs(target - output)),分母用 np.clip(np.sum(mask), 1e-8, np.inf) 防零除。

Failure Modes, Edge Cases & Concurrency

  • .pkl 与 .h5 失配:ScoreCalculator 以 hdf5 文件名定位 pkl(get_filename)。若 infer_prob 与 calculate_metrics 的四元组(model_type / augmentation / dataset / split)任一不同,会得到 FileNotFoundError;这是两阶段流水线最主要的运维坑点。
  • 空踏板参考:len(ref_pedal_on_off_pairs) == 0 时整段踏板指标(含 pedal_frame_*)被跳过,不写入 return_dict;聚合端用 if key in e.keys() 过滤,均值只对"有踏板的曲目"计算——即 pedal_f1 的口径是"踏板曲目子集平均"。
  • 帧数不齐:预测与真值帧数可能差 1(整曲长度取整),双向截断取公共长度;不做任何插值。
  • 阈值恰等:np.sign 产生的 0.5 被显式置 0(偏向保守/低召回),帧级与踏板帧级路径都做了同样处理。
  • 踏板帧级阈值固定 0.5:与音符帧级(可调 frame_threshold)不同,pedal_frame_output 二值化阈值硬编码 0.5,不参与阈值搜索。
  • 并发模型:阶段 2 是进程级并行(ProcessPoolExecutor),无共享状态,calculate_score_per_song 只读自身 pickle,天然无锁。阶段 1 是单进程串行循环(GPU 推理),逐曲打印进度。
  • debug 开关:metrics() 里 debug = False 时直接并行执行;置 True 则串行调用 calculate_score_per_song 并打印曲目路径,便于排查单曲问题(list_args[0:] 保持全集)。
  • __call__ 键名不一致:ScoreCalculator.__call__ 返回 np.mean(stats_dict['f1']),而 metrics() 产出的键是 note_f1 等;把它直接当黑盒优化目标会得到 KeyError,需自行适配。

Performance & Operational Notes

  • 为什么两阶段:一次 infer_prob(GPU)产出概率后,阈值扫描、消融(velocity/pedal/frame 开关)、跨后处理器对比(regression vs onsets_frames)都能在 CPU 上重复进行,无需重跑模型——这是整条流水线最重要的性能设计。
  • 逐曲并行的收益:mir_eval 的区间匹配在长曲(MAESTRO 单曲可达 10+ 分钟、上万音符)上是 CPU 密集操作,进程池按曲粒度并行可近似线性扩展。
  • I/O 布局:每曲一个 .pkl,天然支持按曲并行读取与断点续跑;重命名/移动 probs 目录即失效。
  • 内存:infer_prob 会把整曲概率 + 全部参考标注驻留内存后一次性 pickle;超长曲目内存占用主要来自 (frames, 88) 级别的各输出头。

Extension Points

  • 新指标:在 calculate_score_per_song 末尾向 return_dict 追加键即可自动进入并行、聚合与打印链路(聚合端按键过滤,天然兼容部分曲目缺键)。
  • 阈值搜索:ScoreCalculator.__call__(params) 已具备"参数向量 → 标量目标"形态,可直接接入 hyperopt / optuna 之类的搜索器(注意修正其内部读取的键名)。
  • 对齐其他基线系统:新增后处理器只需在 calculate_score_per_song 的 post_processor_type 分支中注册,并保证实现 output_dict_to_note_pedal_arrays(output_dict) -> (est_on_off_note_vels, est_pedal_on_offs) 契约。
  • 新数据集:CLI 已将 --dataset 约束为 {maestro, maps};只要 utils/features.py 打包出的 .h5 含 split 属性、waveform、midi_event、midi_event_time,评测链路即可复用(打包细节见数据准备相关页面)。
  • pytorch/calculate_score_for_paper.py —— 论文指标计算入口(infer_prob / calculate_metrics / ScoreCalculator)
  • pytorch/evaluate.py —— 训练期 SegmentEvaluator 与 mae
  • runme.sh —— MAESTRO/MAPS 基准评测命令
  • README.md —— MAESTRO V2.0.0 数据集与工作区目录结构说明
  • 推理引擎、后处理器与数据打包的实现细节分别属于推理、后处理与数据准备等兄弟页面。