Show HN: Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO)

Posted by popopanda 11 hours ago

Counter19Comment0OpenOriginal

Comments

Comment by popopanda 11 hours ago

[flagged]