Open to workNew YorkGet in touch

talks · 2026

Taking production AI to the stage.

In 2026 I took the hard part of applied AI to two New York stages - not the demo, but making it run reliably, cheaply, and at scale.

  • Scaling Production RAG Systems with Kubernetes

    Lightning talkKCD New York 2026

    How retrieval-augmented generation earns its uptime: autoscaling vLLM, vector search, and observability on Kubernetes so a RAG system stays fast and affordable when traffic stops being a benchmark and starts being real.

  • Real-Time AI Pipelines at Scale: Embedding LLMs into Apache Beam

    Technical sessionApache Beam Summit 2026

    Bringing LLM inference and RAG inside the pipeline itself - embedding models directly into Apache Beam transforms so streaming data can reason in real time, delivering low-latency intelligence on high-velocity data without ever leaving the stream.

  • Session To Be Announced

    Conference sessionBlockchain Week UNGA Edition 2026

    I will be speaking at Blockchain Week - UNGA Edition 2026, the ten-day independent industry gathering held in New York alongside the UN General Assembly (September 10–19), with tracks spanning Bitcoin, AI agents, energy, and the space economy. The conference will announce the session title and slot; the talk sits where my work does - production AI systems, and what it takes to run them reliably once the demo is over.

Hiring for AI or ML?

I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.