Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up

All HF Hub posts

SeaWolf-AI 
posted an update 1 day ago
view post
Post
3290
🧠 We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything.

Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.

⚙️ How it works
🔹 It makes its call in a single forward pass.
🔹 Zero generated tokens, and no decoding loop.
🔹 That keeps latency and cost far below what a generative model needs.

🎯 What it judges
🔹 It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score).
🔹 For each one it hands back a calibrated confidence, not just an answer.

📊 How well calibrated (measured)
🔹 KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens.
🔹 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors.
🔹 By type: noul 0.847, choice 0.723, score 0.675.
🔹 None of the benchmark's train split went into it. It is pure zero-shot.

🚀 Where it fits
🔹 Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.

🏆 It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).

🔗 Links
Model: FINAL-Bench/Darwin-27B-ZTC
Leaderboard: LocalLLaMA/typed-decisions

Curious to hear what you make of the single-pass, no-generation approach. 🙌
  • 2 replies
·
danielhanchen 
posted an update 1 day ago
raincandy-u 
posted an update 2 days ago
view post
Post
3670
20K parameters can tell a story. 🚀

🤗 We trained a ~20k-parameter Transformer that can actually write stories!

raincandy-u/MacroStories

→ ~50× smaller than the 1M-parameter TinyStories model
→ ~3,000× smaller than AlexNet
→ 81 KB in FP32

yayyy the whole model. ૮ ˶ᵔ ᵕ ᵔ˶ ა

She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block — recurrently applied 4 times with shared weights.

Despite having only 19,969 parameters, she can maintain a simple narrative across 100–300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.

She runs extremely fast on CPU — no GPU required. The entire model is tiny enough to load almost instantly! ☺️
  • 3 replies
·
SeaWolf-AI 
posted an update about 11 hours ago
view post
Post
1119
🏆 Darwin-27B-ZTC-v2 just took #1 on the System One Mosaic Benchmark (S1MB).

S1MB compares 102 models across 137 specialized benchmarks, in three task types: Noul (assess a condition), Choice (select an option), Score (rate on a scale). Ranking is by overall Borda score.

📊 Top of the board
🥇 Darwin ZTC v2 (FINAL-Bench) 89.58
🥈 OpenJev-27B 87.50
🥉 AutoJev-27B 87.07
4️⃣ Eikos 27B 85.43
5️⃣ Jev 1.13 85.05

🔎 Ranks 2 to 5 are all the JEV family (TypeSafe AI's System One model, from ex-OpenAI researchers). S1MB exists to compare these System One judges, so leading it is the headline.

⚙️ Why a zero-token judge wins here
🔹 It does not generate. It reads the input and typed questions and returns a calibrated distribution in a single forward pass.
🔹 Zero generated tokens, no decoding loop, so latency and cost stay low.
🔹 Holds up out of distribution too: General Noul 96.00, General Choice 99.34.

It is also #1 on the typed-decisions leaderboard (0.743, zero-shot). Same message from both: a deterministic, calibrated judge at one forward pass per call.

🔗 Model: FINAL-Bench/Darwin-27B-ZTC-v2
🔗 Leaderboard: hotchpotch/S1MB-leaderboard

Standings move as new models are added. Numbers reflect the board at the time of writing. 🙌
tardellirs 
posted an update 2 days ago
view post
Post
3735
Robotics is now the second most downloaded dataset category on the Hub, after text generation.

Robotics datasets got 13.7M downloads in September, ahead of text classification and question answering. Two years ago the category ranked 23rd. One in 7 new datasets is now robotics, mostly LeRobot recordings: typically a few dozen demos, about half of them on low-cost SO-100/SO-101 arms.

I found this after adding datasets and Spaces to Model Pulse, which rebuilds daily history from @cfahlgren1 's hub-stats snapshots. Two more findings:

- In 2022, 36% of authors who list training data cited classic NLP sets like IMDb, SQuAD and GLUE. In 2026 it's 2.4%. Reasoning traces distilled from models like DeepSeek-R1 and Claude are now the most cited kind.
- In October 2025, 122K Spaces were created, 71K of them websites built with DeepSite. That's about 6x the monthly pace of late 2024, while likes given per month fell from about 35K to about 20K.

New in the app: a page for every dataset, with daily downloads and the models trained on it (636 list FineWeb), a page for every Space, and rankings for both.

Thank you to everyone who liked Model Pulse this week: it made Spaces of the Week and is #7 on trending. Thanks also to @dipankarsarkar , whose comments on the last post fixed three data issues. If a number looks wrong, tell me.

Spaces: tardellirs/model-pulse
  • 6 replies
·
Parveshiiii 
posted an update 3 days ago
view post
Post
3751
Most deepfake audio detectors are quietly cheating.

They don’t really listen to the speech — they just look at how long the embedding vector is. Once they figure that out, accuracy looks great on paper and falls apart in the wild.

AIRealNet-Audio was built to stop that shortcut.
It forces every feature onto the unit hypersphere (twice) so the model can only use direction, not magnitude. Trained on speech from 100+ different TTS and voice-cloning systems, plus real human recordings under heavy compression and noise.

The result is a detector that actually has to learn the artifacts instead of gaming the feature space.

Model: Modotte/AIRealNet-Audio
prithivMLmods 
posted an update 2 days ago
view post
Post
3160
OneDecision-VisionGuard-Demo is now available on Hugging Face Spaces!

🤗 Space: prithivMLmods/OneDecision-VisionGuard-Demo

This demo showcases the OneDecision-VisionGuard family of multimodal image classification models for detecting NSFW and other sensitive visual content, with structured JSON reasoning, improved accuracy, and better handling of edge cases such as sensitive imagery, uncensored analysis, scene descriptions, and classification reasoning.

📦 Models: 27B, 9B, 4B — prithivMLmods/OneDecision-VisionGuard-27B-SFT, prithivMLmods/OneDecision-VisionGuard-9B-SFT, prithivMLmods/OneDecision-VisionGuard-4B-SFT

↗️ Collection: https://huggingface.co/collections/prithivMLmods/onedecision-visionguard

To learn more, visit the app page or the respective model pages.
DedeProGames 
posted an update 3 days ago
view post
Post
3946
Im working on a 23M ASR model, trained on 100k hours of audio
  • 4 replies
·
danielhanchen 
posted an update 3 days ago
Parveshiiii 
posted an update about 5 hours ago
view post
Post
38
Most NSFW classifiers break the second an image touches the internet.

They look great on pristine benchmarks, but in the wild, every social platform aggressively recompresses, downsamples, and degrades images.
The moment JPEG or WebP compression artifacts show up, confidence collapses and false positives spike.

SafeScan was built to survive actual platform pipelines.
Trained on 34,000 images under almost every major social media compression profile using a Vision Transformer backbone (google/vit-base-patch16-224). Instead of blunt binary filtering, it breaks decisions down across 5 clear categories:

• safe
• drawing
• sexy
• hentai
• porn

The result is a moderation model that actually generalizes to real-world internet feeds instead of fragile, uncompressed datasets.
Open-weight and available on Hugging Face:

Model: Parveshiiii/SafeScan