Open to workNew YorkGet in touch

projects

Twelve systems, eleven in the open.

An AI harness for VS Code, an open-weight decision model, on-device LLM inference, RAG pipelines, multi-agent systems and the MLOps that keeps them honest - each with a write-up of the problem, how it works, and why it exists.

  • Omnitron - AI coding agents in your own workspace

    TypeScriptAgentsElectron

    My build of DeepSeek Harness, the open-source platform for running AI coding agents in your own workspace: one plugin core behind a web UI and a desktop app for Linux, macOS and Windows, released as Omnitron 1.0 with its own identity, interface and operational guards.

  • Tron-1B - A fast, open-weight decision model

    PythonPyTorchHugging Face

    An open-weight 1.1B decision model with its own runtime and decision head. It answers choice, rating and yes/no questions with calibrated probabilities in about 17 ms, and beats Jev 1.13.0 on all four of its published benchmarks (average 90.1 vs 74.7). It improves itself nightly from a larger model's labels.

  • Wayne - AI harness for VS Code

    VS CodeAgentsMCP

    An AI harness living in the IDE - it plans before it codes, reads the codebase and follows call sites, runs commands and reacts to failures, and presents every change as a diff you approve. Extends through MCP. Live on the VS Code Marketplace.

  • PrometheusAI - Offline LLM assistant for Android

    KotlinLLaMAGGUF

    On-device inference with zero cloud dependency. Runs LLaMA 3.1, IBM Granite, and Mistral via 4-bit GGUF at ~8 tok/s, with on-device RAG over your own documents.

  • tokenash - Local-first context compression

    PythonLLMCompression

    Context compression for LLM applications that runs locally - 50–90% fewer tokens for the same answers, with nothing shipped to a third-party service to get there.

  • HLST - Long-context LLMs on edge

    PythonResearchState-Space

    Reference code for my TechRxiv preprint: hierarchical latent state-space recurrence that democratises long-context LLMs on edge devices, as a transformer alternative.

  • KCD New York 2026 - Production RAG on Kubernetes

    PythonKubernetesvLLM

    Demo from my KCD New York 2026 talk - deploying, autoscaling, and observing production RAG pipelines on Kubernetes with vLLM and vector search.

  • Apache Beam Summit 2026 - Real-time AI pipelines

    PythonApache BeamLLM

    Embedding LLMs and RAG directly into Apache Beam transforms for low-latency inference on high-velocity data streams.

  • Relay - Open-source outbound agent

    PythonAgentsAutomation

    An open-source AI agent for outbound leads - the research, qualification, and follow-up loop handled by a tool-using agent rather than a sequence of manual steps.

  • Franola - Open-source AI note-taking

    RustLLMNotes

    An open-source AI note-taking app in the Granola mould, written in Rust - meeting notes captured and structured by a model, without the notes leaving a proprietary product.

  • GateKeep - Agentic personal finance

    JavaScriptAgentsFinance

    An agentic-AI finance tracker that ingests transactions and reasons over them with tool-using agents to surface spend insights.

  • AIORNOR - AI-or-human text detection

    PythonNLPPerplexity

    Detects whether a piece of text is AI-generated or human-written using a perplexity-based metric, served through a lightweight web app.

Hiring for AI or ML?

I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.