I tried to build a specialized 8B security model. The most interesting result was that I had to reject it

I’ve been experimenting with a question that I think goes beyond cybersecurity:

What if an AI model is allowed to reason, but is never allowed to be the source of truth?

I built an open-source research prototype called XSS Specialist, using XSS vulnerability analysis as the test environment.

The original goal was relatively straightforward: take a small open 8B model and see how far specialization could push it using retrieval, LoRA, gated continual learning, and adversarial evaluation.

The unexpected part was what happened next.

The specialist became significantly better at several parts of the task, but repeatedly failed an adversarial near-miss test. Small changes to security-relevant identifiers could still cause incorrect safety judgments.

Fine-tuning didn’t eliminate it.

Retrieval didn’t eliminate it.

Multiple specialist iterations didn’t eliminate it.

So instead of lowering the benchmark, the promotion gate rejected every model candidate.

That changed the direction of the experiment.

I redesigned the system around a different assumption:

The model can reason. The model can propose. But the model cannot establish truth.

For live assessment, confirmation authority was moved outside the model to an independent execution layer. A potential XSS finding only becomes CONFIRMED when a real headless browser actually executes a benign sentinel.

On the frozen 102-case live benchmark, the resulting system achieved:

  • 94.1% precision

  • 100% recall

  • 0% false-negative rate

  • 0.00 near-miss sanitizer leakage at the execution boundary

  • 13/13 predefined acceptance criteria passed

The benchmark is synthetic/local and the project is still a research prototype, so I don’t consider these numbers evidence of production readiness.

The result I find more interesting is architectural:

Maybe reliable AI systems shouldn’t depend on making the model trustworthy enough.

Maybe they should be designed to remain reliable when the model is wrong.

In this experiment:

The model failed. The architecture didn’t.

I’ve open-sourced the code, architecture, benchmarks, evaluation methodology, failures, and negative results:

I’d particularly appreciate feedback from people working on small-model specialization, agents, model evaluation, verification, or AI security.

I’m curious: Where else could this separation between model reasoning and external verification be useful?