Open to workNew YorkGet in touch

flagship · open weights · project

Tron-1B

A fast, open-weight decision model

Built with
Python · PyTorch
Source
huggingface.co
Author
Samir Sengupta

A 1.1-billion-parameter model that answers typed questions about any text (pick an option, rate on a scale, or yes/no) with calibrated probabilities in about 17 ms. It beats Jev 1.13.0 on all four of Jev's published benchmarks.

Bar chart: Tron-1B scores 90.1 on the Decision Index, the average of Jev's four published benchmarks; Jev 1.13.0 scores 74.7.

Fine print Tron trained on the training splits of Banking77, AG News and emotion, and was scored on their separate test splits. Jev's Banking77 result used 72 labels; Tron's used all 77. Jev's latency was measured through its hosted API, so it includes network time.

The problem

AI apps and agents make thousands of small decisions: where should this ticket go, is this a prompt-injection attempt, what does the user want, is this step done? Sending each one to a large LLM is slow, expensive and inconsistent. Tron handles them in a single forward pass. It returns a probability you can set thresholds on, and it can decline to answer when it's unsure.

What I built

  • troncore runtime and decision head, my own inference engine and model head instead of an off-the-shelf classifier. It reads the question, the input and every option together, pools each option into a vector, and compares all the options with each other through a small attention layer before scoring them. That's what separates close labels like "card not arrived" and "card delivery estimate". Probabilities are calibrated separately for each question type and option count. It handles more than 128 options by scoring them in rounds, and reads long inputs in overlapping windows.
  • Training pipeline: sampling that keeps large datasets from crowding out small ones, a loss for rating questions that accounts for how far off a guess is, smaller learning rates for lower encoder layers, batches packed to a token budget, best-checkpoint selection and its own checkpoint format.
  • Data: 1.19 million typed decisions built from 368 public datasets, covering intent, topic, sentiment, safety, prompt injection, routing, reasoning, preference and business workflows.
  • Training: about 13 hours on a single NVIDIA DGX Spark (GB10), starting from the MIT-licensed ettin-encoder-1b.

The results

Bar charts per benchmark, Tron-1B against Jev 1.13.0: Banking77 94.0% vs 87.0%, AG News 93.9% vs 91.0%, typed business decisions 79.6% vs 72.7%, DAIR emotion 92.9% vs 48.0%, and latency per decision 16.7 ms vs 236-276 ms.
Tron-1B against Jev 1.13.0 on Jev's published benchmarks
BenchmarkTron-1BJev 1.13.0
Banking77 (77 intents)94.0%87.0%
AG News (topic)93.9%91.0%
Typed business decisions79.6%72.7%
DAIR emotion92.9%48.0%
Latency per decision (p50)16.7 ms236–276 ms
Decision Index (average)90.174.7
On tasks it never trained on
TaskAccuracy
IMDB sentiment94.7%
Prompt-injection detection90.5%
MASSIVE intent (59 labels)88.4%
XNLI entailment86.7%
BoolQ79.2%

Fine print Tron trained on the training splits of Banking77, AG News and emotion, and was scored on their separate test splits. Jev's Banking77 result used 72 labels; Tron's used all 77. Jev's latency was measured through its hosted API, so it includes network time.

Engineering around the model

  • Private inference API serving Tron-1B alongside specialised models, with API-key authentication.
  • Batch endpoint that scores thousands of inputs in one request: 11 times faster than one at a time, at about 440 decisions a second.
  • Nightly self-improvement loop: logs real requests, picks the ones Tron was least sure about, has a local 27B LLM label them as a teacher, and fine-tunes Tron on those labels mixed with its original training data. The new version goes live only if it beats the current one on held-back questions and holds every benchmark; otherwise it's rejected.
  • Open-weight release: Hugging Face model card, benchmark charts, the runtime as a pip-installable wheel, and a Colab quickstart notebook.

Tech and licences

Tech: Python, PyTorch, Hugging Face Transformers and Hub, safetensors, ModernBERT-style encoders, NVIDIA DGX Spark, Tailscale, systemd, SGLang, Qwen3.8-27B (teacher).

Licences: weights CC BY-NC 4.0; troncore runtime Apache-2.0.

More work

Hiring for AI or ML?

I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.