Connecting to the evidence…
Watch Jude learn.
Teach difficult things. Test new examples without the teacher. Keep only improvements that last.
—new skills kept todayChecking evidence
—spent on today's runsCost not measured yet
—unseen test gainNo teacher-off comparison yet
What is happening
- Waiting for measured events.
Does Jude need a bigger brain yet?
No evidence yet.
A flat score is a reason to investigate. It does not prove that a bigger model is better.
What would prove it?
Compare bases on the same unseen tests, retained old skills and total cost. Check data quality, training settings and forgetting before blaming model size.
Examples never count as measured results.
Chosen starting modelQwen3.8-27B
TeacherK3 · selected, connection not yet verified
What do the numbers mean?
A kept skill needs a new unseen test, the teacher off, old skills retained, and a fresh reload. Test points describe a named benchmark, not general intelligence. Missing cost records keep £ yields blank. Today's costs include failed and inconclusive runs; repeated skills count once.