Amod Sahabandu

Amod Sandeepa Sahabandu

I turn frontier AI into production systems.

I work where new AI capabilities meet real enterprise constraints. I find the useful edge, shape the problem, build the system, and stay with the work until it operates reliably in production.

  1. 01Assess
  2. 02Shape
  3. 03Build
  4. 04Operate

From capability to operation.

The work is not finished when a model responds. It is finished when the system behaves reliably under real constraints.

  1. Assess

    01

    Find the useful edge before committing to a system.

    • Test emerging capabilities against the real problem, not a demo-shaped version of it.
    • Define success, failure modes, evidence requirements, and operating constraints early.
  2. Shape

    02

    Turn an ambiguous opportunity into an engineering problem.

    • Choose model, retrieval, tool, interface, and approval boundaries deliberately.
    • Make cost, latency, security, privacy, and reliability trade-offs explicit.
  3. Build

    03

    Engineer the system and the evidence around it.

    • Build agentic, retrieval, document, model-adaptation, and evaluation systems.
    • Preserve provenance, structured outputs, observability, and reproducibility.
  4. Operate

    04

    Stay with the work until it holds up in production.

    • Own deployment, monitoring, escalation, failure handling, and iteration.
    • Treat production behavior—not prototype quality—as the final test.

Systems I have taken from problem definition to production.

01 / Flagship system

Reflective Prompt Optimization

A production optimization system that evolves prompts against measured behavior instead of intuition.

Read the RealAIzation case study

Manual prompt iteration does not scale, hides evaluation blind spots, and makes regressions difficult to explain.

Separated candidate, judge, and mutator roles combine deterministic checks, semantic evaluation, feedback-driven mutation, and Pareto selection.

  • Use minibatches and accept-if-better policies to control evaluation cost and regressions.
  • Track candidate lineage, model cycling, stopping criteria, and instruction drift.
  • Keep evaluation traces reproducible enough to explain why a prompt advanced.
  • Runs in production for a healthcare services provider.
  • Maintains evaluation and lineage traces for every accepted candidate.
02

Provenance-first Document Intelligence

Extraction and retrieval systems designed for documents whose answers must survive scrutiny.

Dense reports, scans, tax documents, and credit agreements fail conventional extraction when structure and image quality vary.
Denoising, OCR preparation, extraction, retrieval, and evaluation are joined by source-region and cell-level provenance.
  • Preserve the exact source location behind each extracted value.
  • Evaluate document quality and extraction behavior as separate failure surfaces.
  • Prefer traceable abstention over unsupported completion.
  • Shipped an end-to-end denoising platform for messy scanned documents.
  • Designed lineage from spreadsheet cells and extracted values back to source material.
03

Automated Translation Quality Assessment

A fully automated production system using methods validated against human judgments to surface meaning changes, rank language risk, and preserve the evidence behind every result.

Aggregate quality scores could not show which languages were failing, where meaning changed, or what required human review.
A staged evaluation and correction pipeline combines MetricX scoring, distilled XCOMET screening, GEMBA v2, and tagged span annotation. A review console ranks languages by concern while preserving the segment-level evidence behind every result.
  • Benchmark newly published WMT25 methods against the existing evaluation stack before choosing what to productionize.
  • Combine quality, severity, and span-level signals instead of treating one metric as ground truth.
  • Carry evaluation through correction, re-checking, and an executive language view without losing the underlying evidence.
  • Moved from the late-November WMT25 proceedings to an executive-ready demonstration in January 2026.
  • Own the production ML architecture and system end to end as its sole ML engineer.

I publish open work that practitioners can use.

Mental Health Counseling Conversations

An open counselling-conversation dataset released for machine-learning research and downstream model development.

100k+
2023
10.57967/hf/1581
RAIL-D
View on Hugging Face

Responsibility moved faster than the calendar.

  1. RealAIzation

    AI Consultant → Head of Engineering

    Sep 2024 — Present

    Lead engineering and AI delivery across enterprise and private-equity portfolio engagements.
    Own technical diligence, system architecture, implementation, evaluation, and production delivery across agentic, retrieval, and document-intelligence work.
  2. Arcadea Group

    Consultant AI Engineer

    Jul 2025 — Jan 2026

    Advised on AI strategy for vertical-software acquisitions and pipeline improvement.
    Connected technical feasibility with the operating context of acquired software businesses.
  3. Altrium

    AI/ML Engineer

    Jun 2024 — Aug 2025

    Worked end-to-end on generative-AI systems for PlanYear and related product initiatives.
    Shipped production AI for a platform serving US healthcare-insurance brokerages.
  4. Insighture

    Associate Machine Learning Engineer

    Aug 2023 — May 2024

    Designed and deployed generative-AI features for SkyU.io.
    Moved early AI capabilities into a DevOps-automation product.
BSc (Hons) Electronics & Telecommunications Engineering · Sri Lanka Technological Campus · 2019 — 2023

Notes from the work.

View all writing

The engineering is serious. The person remains broader than the work.

Outside production AI, I spend time with audio, aviation, astrophysics, philosophy, nuclear science, and global politics.