The 500M instruction model aligned with DPO on 4,000 preference pairs. Ask a question or give an instruction.
The 500M instruction model aligned with DPO (direct preference optimization, beta=0.1); the frozen reference is the instruction model, on 4,000 AI-feedback preference pairs spanning closed-book QA and instruction-following failure modes (wrong figures, invented citations, broken format constraints, ignored instructions), prompts held out of every SFT set. Lineage: base → QA SFT → instruction SFT → DPO v2.
Served scale-to-zero on Modal, so the first request may take ~20–60s while the model wakes.