Open to work
Hire an AI/ML engineer in New York.
I am Samir Sengupta — an AI/ML Engineer and Data Scientist in New York who takes a model from a framed problem to a served, monitored endpoint. Production LLM, RAG, and multi-agent systems in Python, engineered onto Kubernetes with the evaluation infrastructure that keeps them honest.
Most teams hiring for AI right now have the same problem: a prototype that works in a notebook and nobody who can make it survive contact with real traffic. That gap is the work I do.
Across 3+ years I have framed problems as a data scientist, built the models, and engineered them into reliable software — Kafka pipelines and FastAPI services on Kubernetes at sub-100ms latency, retrieval systems that stay fast when the corpus stops being a demo, and predictive models that cut churn and lifted conversion for the teams that shipped them.
What I build
Production LLM & RAG systems
Retrieval that holds up under load: hybrid search over a vector index, cross-encoder reranking, vLLM behind FastAPI on Kubernetes with continuous batching and KV-cache reuse. Autoscaling on real signals, not CPU guesswork.
Agentic AI & multi-agent orchestration
Agents as state machines rather than prompt chains — explicit nodes for retrieval, tool calls, and self-checks, with retries where failure is expected. Roughly 95% task success on the enrichment system I built this way.
Machine learning & data science
Forecasting, anomaly detection, churn and propensity modelling, and the A/B testing programme that proves any of it moved a number. Rigorous evaluation before the dashboard, not after.
MLOps, serving & data engineering
The half that decides whether a model is a science project or a system: streaming pipelines, CI/CD, model monitoring, tracing on every decision, and drift alerts that fire before a customer writes in.
How we can work together
- Full-time roles. AI/ML Engineer, Machine Learning Engineer, Data Scientist, or Python Engineer. New York and the wider US, on-site, hybrid, or remote.
- Contract & consulting. Shipping a stalled LLM or RAG prototype into production, or standing up the serving and evaluation infrastructure a team is missing.
- Research collaboration. Efficient long-context architectures and edge inference — the line of work behind my TechRxiv preprint and both 2026 conference talks.
Questions I get asked
What kind of AI/ML engineering roles are you open to?
AI/ML Engineering, Machine Learning Engineering, and Data Science roles — full-time, contract, or consulting. I work across the whole lifecycle: framing the problem, building the model, and engineering it into production software. I am based in New York and available to teams in the US and remotely worldwide.
What can you build?
Production LLM and RAG systems, multi-agent orchestration, fine-tuning and model serving, plus the classical ML behind forecasting, anomaly detection, churn prediction and A/B testing. On the infrastructure side: FastAPI services on Kubernetes, Kafka and Apache Beam pipelines, CI/CD, evaluation suites and model observability.
What have you shipped in production?
A Kafka pipeline processing around 500,000 events a day, ML predictions served on Kubernetes at sub-100ms p50 latency, an NLP ticket-routing system that cut manual triage by about 40%, churn models that reduced monthly attrition by about 25%, and an A/B testing programme that lifted conversion by about 32%.
What is your background?
Over 3 years across data science and software engineering, an MS in Data Science (GPA 3.9), a TechRxiv research preprint on efficient long-context LLMs, speaking slots at KCD New York 2026 and the Apache Beam Summit 2026, and 10+ certifications from AWS, Google, Microsoft, NVIDIA, IBM and Stanford.
How do I get in touch?
Email [email protected] or call +1 (551) 359-1228. Email is the fastest route — it goes straight to my inbox.
If any of that maps onto what you are hiring for, email me. I read every message and reply within a day.
Recent work
Let’s ship AI that survives production.
I’m the author of a research preprint on efficient long-context LLMs, spoke at KCD New York and the Apache Beam Summit in 2026, and hold an MS in Data Science (GPA 3.9) plus industry certifications from AWS, Google, NVIDIA, and IBM.