flagship · open weights · project
Tron-1B
A fast, open-weight decision model
A 1.1-billion-parameter model that answers typed questions about any text (pick an option, rate on a scale, or yes/no) with calibrated probabilities in about 17 ms. It beats Jev 1.13.0 on all four of Jev's published benchmarks.

Fine print Tron trained on the training splits of Banking77, AG News and emotion, and was scored on their separate test splits. Jev's Banking77 result used 72 labels; Tron's used all 77. Jev's latency was measured through its hosted API, so it includes network time.
The problem
AI apps and agents make thousands of small decisions: where should this ticket go, is this a prompt-injection attempt, what does the user want, is this step done? Sending each one to a large LLM is slow, expensive and inconsistent. Tron handles them in a single forward pass. It returns a probability you can set thresholds on, and it can decline to answer when it's unsure.
What I built
- troncore runtime and decision head, my own inference engine and model head instead of an off-the-shelf classifier. It reads the question, the input and every option together, pools each option into a vector, and compares all the options with each other through a small attention layer before scoring them. That's what separates close labels like "card not arrived" and "card delivery estimate". Probabilities are calibrated separately for each question type and option count. It handles more than 128 options by scoring them in rounds, and reads long inputs in overlapping windows.
- Training pipeline: sampling that keeps large datasets from crowding out small ones, a loss for rating questions that accounts for how far off a guess is, smaller learning rates for lower encoder layers, batches packed to a token budget, best-checkpoint selection and its own checkpoint format.
- Data: 1.19 million typed decisions built from 368 public datasets, covering intent, topic, sentiment, safety, prompt injection, routing, reasoning, preference and business workflows.
- Training: about 13 hours on a single NVIDIA DGX Spark (GB10), starting from the MIT-licensed ettin-encoder-1b.
The results

| Benchmark | Tron-1B | Jev 1.13.0 |
|---|---|---|
| Banking77 (77 intents) | 94.0% | 87.0% |
| AG News (topic) | 93.9% | 91.0% |
| Typed business decisions | 79.6% | 72.7% |
| DAIR emotion | 92.9% | 48.0% |
| Latency per decision (p50) | 16.7 ms | 236–276 ms |
| Decision Index (average) | 90.1 | 74.7 |
| Task | Accuracy |
|---|---|
| IMDB sentiment | 94.7% |
| Prompt-injection detection | 90.5% |
| MASSIVE intent (59 labels) | 88.4% |
| XNLI entailment | 86.7% |
| BoolQ | 79.2% |
Fine print Tron trained on the training splits of Banking77, AG News and emotion, and was scored on their separate test splits. Jev's Banking77 result used 72 labels; Tron's used all 77. Jev's latency was measured through its hosted API, so it includes network time.
Engineering around the model
- Private inference API serving Tron-1B alongside specialised models, with API-key authentication.
- Batch endpoint that scores thousands of inputs in one request: 11 times faster than one at a time, at about 440 decisions a second.
- Nightly self-improvement loop: logs real requests, picks the ones Tron was least sure about, has a local 27B LLM label them as a teacher, and fine-tunes Tron on those labels mixed with its original training data. The new version goes live only if it beats the current one on held-back questions and holds every benchmark; otherwise it's rejected.
- Open-weight release: Hugging Face model card, benchmark charts, the runtime as a pip-installable wheel, and a Colab quickstart notebook.
Tech and licences
Tech: Python, PyTorch, Hugging Face Transformers and Hub, safetensors, ModernBERT-style encoders, NVIDIA DGX Spark, Tailscale, systemd, SGLang, Qwen3.8-27B (teacher).
Licences: weights CC BY-NC 4.0; troncore runtime Apache-2.0.
More work
Hiring for AI or ML?
I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.
