Apriel-H1: Towards Efficient Enterprise Reasoning Models

Posted by guiriduro 3 hours ago

Counter1Comment1OpenOriginal

Comments

Comment by guiriduro 3 hours ago

Apriel-H1-15b-Thinker-SFT uses incremental distillation from Apriel-Nemotron-15B-Thinker, selectively replacing less critical attention layers with linear Mamba blocks to reduce computational complexity while preserving reasoning quality.