Open to work

Hire an AI/ML engineer in New York.

I am Samir Sengupta — an AI/ML Engineer and Data Scientist in New York who takes a model from a framed problem to a served, monitored endpoint. Production LLM, RAG, and multi-agent systems in Python, engineered onto Kubernetes with the evaluation infrastructure that keeps them honest.

Based in
New York, USA
Status
Open to work
Experience
3+ years
Reply time
Within 24 hours
500K+
events / day through a Kafka pipeline in production
<100ms
p50 ML prediction latency, FastAPI on Kubernetes
~95%
task success on a multi-agent enrichment system
~40%
manual triage removed by NLP ticket routing
~32%
conversion lift from an A/B testing programme
10+
certifications — AWS, Google, Microsoft, NVIDIA, IBM

Most teams hiring for AI right now have the same problem: a prototype that works in a notebook and nobody who can make it survive contact with real traffic. That gap is the work I do.

Across 3+ years I have framed problems as a data scientist, built the models, and engineered them into reliable software — Kafka pipelines and FastAPI services on Kubernetes at sub-100ms latency, retrieval systems that stay fast when the corpus stops being a demo, and predictive models that cut churn and lifted conversion for the teams that shipped them.

What I build

Production LLM & RAG systems

Retrieval that holds up under load: hybrid search over a vector index, cross-encoder reranking, vLLM behind FastAPI on Kubernetes with continuous batching and KV-cache reuse. Autoscaling on real signals, not CPU guesswork.

LangGraphvLLMFAISSPineconeAWS BedrockAzure OpenAI

Agentic AI & multi-agent orchestration

Agents as state machines rather than prompt chains — explicit nodes for retrieval, tool calls, and self-checks, with retries where failure is expected. Roughly 95% task success on the enrichment system I built this way.

LangGraphLangChainCrewAIGuardrails.aiDSPy

Machine learning & data science

Forecasting, anomaly detection, churn and propensity modelling, and the A/B testing programme that proves any of it moved a number. Rigorous evaluation before the dashboard, not after.

PyTorchscikit-learnXGBoostPandasA/B Testing

MLOps, serving & data engineering

The half that decides whether a model is a science project or a system: streaming pipelines, CI/CD, model monitoring, tracing on every decision, and drift alerts that fire before a customer writes in.

KubernetesDockerKafkaApache BeamMLflowTerraform

How we can work together

  • Full-time roles. AI/ML Engineer, Machine Learning Engineer, Data Scientist, or Python Engineer. New York and the wider US, on-site, hybrid, or remote.
  • Contract & consulting. Shipping a stalled LLM or RAG prototype into production, or standing up the serving and evaluation infrastructure a team is missing.
  • Research collaboration. Efficient long-context architectures and edge inference — the line of work behind my TechRxiv preprint and both 2026 conference talks.

Questions I get asked

What kind of AI/ML engineering roles are you open to?

AI/ML Engineering, Machine Learning Engineering, and Data Science roles — full-time, contract, or consulting. I work across the whole lifecycle: framing the problem, building the model, and engineering it into production software. I am based in New York and available to teams in the US and remotely worldwide.

What can you build?

Production LLM and RAG systems, multi-agent orchestration, fine-tuning and model serving, plus the classical ML behind forecasting, anomaly detection, churn prediction and A/B testing. On the infrastructure side: FastAPI services on Kubernetes, Kafka and Apache Beam pipelines, CI/CD, evaluation suites and model observability.

What have you shipped in production?

A Kafka pipeline processing around 500,000 events a day, ML predictions served on Kubernetes at sub-100ms p50 latency, an NLP ticket-routing system that cut manual triage by about 40%, churn models that reduced monthly attrition by about 25%, and an A/B testing programme that lifted conversion by about 32%.

What is your background?

Over 3 years across data science and software engineering, an MS in Data Science (GPA 3.9), a TechRxiv research preprint on efficient long-context LLMs, speaking slots at KCD New York 2026 and the Apache Beam Summit 2026, and 10+ certifications from AWS, Google, Microsoft, NVIDIA, IBM and Stanford.

How do I get in touch?

Email [email protected] or call +1 (551) 359-1228. Email is the fastest route — it goes straight to my inbox.

If any of that maps onto what you are hiring for, email me. I read every message and reply within a day.

Recent work

See the full portfolio

Let’s ship AI that survives production.

I’m the author of a research preprint on efficient long-context LLMs, spoke at KCD New York and the Apache Beam Summit in 2026, and hold an MS in Data Science (GPA 3.9) plus industry certifications from AWS, Google, NVIDIA, and IBM.