Post
3290
🧠 We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything.
Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.
⚙️ How it works
🔹 It makes its call in a single forward pass.
🔹 Zero generated tokens, and no decoding loop.
🔹 That keeps latency and cost far below what a generative model needs.
🎯 What it judges
🔹 It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score).
🔹 For each one it hands back a calibrated confidence, not just an answer.
📊 How well calibrated (measured)
🔹 KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens.
🔹 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors.
🔹 By type: noul 0.847, choice 0.723, score 0.675.
🔹 None of the benchmark's train split went into it. It is pure zero-shot.
🚀 Where it fits
🔹 Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.
🏆 It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).
🔗 Links
Model: FINAL-Bench/Darwin-27B-ZTC
Leaderboard: LocalLLaMA/typed-decisions
Curious to hear what you make of the single-pass, no-generation approach. 🙌
Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.
⚙️ How it works
🔹 It makes its call in a single forward pass.
🔹 Zero generated tokens, and no decoding loop.
🔹 That keeps latency and cost far below what a generative model needs.
🎯 What it judges
🔹 It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score).
🔹 For each one it hands back a calibrated confidence, not just an answer.
📊 How well calibrated (measured)
🔹 KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens.
🔹 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors.
🔹 By type: noul 0.847, choice 0.723, score 0.675.
🔹 None of the benchmark's train split went into it. It is pure zero-shot.
🚀 Where it fits
🔹 Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.
🏆 It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).
🔗 Links
Model: FINAL-Bench/Darwin-27B-ZTC
Leaderboard: LocalLLaMA/typed-decisions
Curious to hear what you make of the single-pass, no-generation approach. 🙌