Vizuara AI Labs · preference-aligned (DPO v2)

SLM‑500M  DPO

The 500M instruction model aligned with DPO on 4,000 preference pairs. Ask a question or give an instruction.

500M
parameters
4,000
pref pairs
0.880
margin acc
β=0.1
dpo
Validation metrics along the 500M lineage
Each perplexity is measured on that stage's own validation set, so read the trend as 'how well the model fits its own stage's data', not as one curve on one dataset. DPO and RLAIF optimize preferences rather than likelihood, so they log preference margin and reward instead of perplexity. Click a stage to open that model.
Base
ppl 7.91
pretrain val
QA SFT
ppl 5.41
QA val
Instruct
ppl 6.69
instruction val
DPO
margin 88.0%
preference val, no ppl
/
RLAIF
reward 6.7→10.1
RM reward, no ppl
RAFT on DPO
ppl 1.95
RAFT val
/
RAFT on RLAIF
ppl 2.01
RAFT val
instruction or question
optional: text to work on (attached as TEXT)
ready
The response will appear here.

What this is DPO v2

The 500M instruction model aligned with DPO (direct preference optimization, beta=0.1); the frozen reference is the instruction model, on 4,000 AI-feedback preference pairs spanning closed-book QA and instruction-following failure modes (wrong figures, invented citations, broken format constraints, ignored instructions), prompts held out of every SFT set. Lineage: base → QA SFT → instruction SFT → DPO v2.

Served scale-to-zero on Modal, so the first request may take ~20–60s while the model wakes.