About

For six years I have built enterprise AI for production, from multi-agent platforms to edge inference, and led the teams that ship it across cloud, on-prem, and air-gapped environments.

I am an Engineering Manager at Iterate.ai, where I lead a team of 12 across frontend, backend, DevOps, and AI/ML, including 2 tech leads and 2 project managers. We build Generate, the company's enterprise AI platform. I joined as a data scientist and moved into the manager role in under two years. Generate has since won AI Product of the Year from both Pinnacle and TMC, and it ships through partnerships with IBM, NetApp, Intel, AMD, and HP.

The work stays hands-on. I hold an MASc in Electrical & Computer Engineering from McMaster University, where my thesis was SciRAG, a retrieval-focused fine-tuning strategy for scientific documents. I have 4 filed patents and peer-reviewed publications with 251 citations.

RAG · Multi-Agent Systems · Agent Runtimes & Memory · Edge & On-Prem Inference · Enterprise AI Security · Team Leadership

Impact at a glance

99.8%
Platform uptime
at peak load
33%
Infrastructure cost
reduction
10K+
Intel AI PC units
shipped
12
Team members led
(scaled from 6)
4
US patents
filed
251
Research citations

Case study: Generate Iterate.ai →

Generate is Iterate.ai's enterprise AI platform. It runs agentic workflows and document intelligence over a company's private data. I have owned its architecture since it was a 5-engineer edge project, and I own it now as a multi-tenant SaaS product that also ships air-gapped into regulated industries. It is the same system in both places, which is most of what makes it hard.

The architecture

Generate runs as event-driven microservices that deploy identically to public cloud, customer-managed on-prem, and edge hardware, and air-gapped installs get the same feature set with no outbound network. That constraint shaped nearly every decision. Model serving is hardware-agnostic across the major accelerator vendors. Low-bit quantization keeps capable models running on consumer-grade machines, so private document search can run entirely on a user's own laptop.

An extraction pipeline handles the documents that defeat naive retrieval: nested tables, embedded visuals, and mixed formats. Accuracy is 94%, and a patent-filed table-extraction method reaches 96% precision on tables. Agent sessions run 30 to 40 minutes and outgrow even a million-token context window. A graph-based memory layer carries user context across them by semantic recall, compacts the working context as it fills, and narrows oversized tool output recursively instead of truncating it.

Agents write and run code, so every session executes inside its own kernel-isolated sandbox and a compromised session cannot reach another tenant's. Enterprise identity resolves through to file-level permissions, which means retrieval is filtered by what each user is cleared to open rather than trimmed after the fact. Administrators bind a model to each AI service, and every binding passes a standardized compatibility suite before it can be used. The platform holds 99.8% uptime at 50,000+ requests/second. An architecture consolidation with connection pooling cut infrastructure cost 33% and improved performance 4×.

What it produced
  • DealsTwo enterprise contracts. I delivered on the engineering commitments behind both.
  • Reach10,000+ Intel AI PC units shipped via the Intel Software Advantage Program, and demoed at the Intel Vision 2024 keynote.
  • PartnerIBM, NetApp, Intel, AMD, HP. The healthcare launch with Hutchinson Regional Healthcare as first customer was featured in IBM's product blog.
  • StorageNetApp ONTAP integration covering snapshot and SnapMirror through consistency groups, ACL-aware access to bound NFS shares, and regional indexing that keeps customer data in-region.
  • IP3 patents filed from the platform work, covering agent creation and routing, AI workflow generation, and LLM fine-tuning.

Experience Full history on LinkedIn →

Engineering Leadership

Engineering Manager

May 2025 to Present

Iterate.ai, Toronto, ON

  • Scaled the team from 6 to 12 across frontend, backend, DevOps, and AI/ML, including 2 tech leads and 2 project managers. I led 6 hires and promoted 2 ICs to Senior Applied AI Engineer. I also mentored 4 engineers and 3 interns, one of whom now owns the agents platform.
  • Led the engineering behind two enterprise contracts for Generate. The product went on to win AI Product of the Year from both Pinnacle and TMC.
  • Own engineering for Generate across cloud, on-prem, and edge (AWS, IBM Cloud, Intel/AMD/NVIDIA), so it ships both as SaaS and as an air-gapped enterprise install.
  • Architected an 8-microservice event-driven platform at 99.8% uptime and 50,000+ requests/second at peak, with multi-terabyte indexing.
  • Built a document intelligence pipeline for tables, visuals, and mixed formats. Optimizing the SQL and vector database integration got it to 94% extraction accuracy and 400% faster responses.
  • Own the cloud and tooling budget: cut infrastructure cost 33% and improved system performance 4× through architectural consolidation, PgBouncer connection pooling, and edge-inference optimization.
  • Launched Generate for healthcare revenue recovery via an IBM reseller partnership with Hutchinson Regional Healthcare, featured in IBM's product blog.

Engineering Manager (Contract)

Sep 2024 to May 2025

Iterate.ai, Remote, Canada

  • Took ownership of a 5-engineer team and defined the roadmap that made Generate Iterate.ai's primary enterprise revenue product.
  • Built an event-driven microservices platform (Kafka, FastAPI, LangGraph) for high-throughput AI workloads, and set up the engineering operating model the full-time team later inherited.
Applied AI & Research

Data Scientist

Nov 2023 to Sep 2024

Iterate.ai, Remote, Canada

  • Led a 5-engineer team building edge-deployed AI for private document search on Intel AI PCs. It shipped to 10,000+ units through the Intel Software Advantage Program.
  • Built an Outlook RAG plugin for email automation (85% time savings, 92% user satisfaction) and presented it at the Intel Vision 2024 keynote.
  • Built a novel table-extraction system (Non-Maximum Suppression + Vision Language Models) at 96% precision; filed US Patent App. 18/590,347.
  • Fine-tuned an LLM with 4/8-bit quantization for a food-ordering chatbot: order accuracy up 40%, unintended responses down 78%.

Earlier ML Engineering

2021 to 2023

ExentAI & Simplyfai, Sri Lanka (Remote)

  • ML Research Engineer, ExentAI: built a newspaper digitization service processing 65,000+ historical documents with multilingual OCR (Tamil/Sinhala/English) on AWS with Docker & Kubernetes. It handled 10TB+ of data at 45% lower memory usage. Also shipped LLM pipelines for sentiment, emotion, and hate-speech detection.
  • Engineering Intern, Simplyfai: built fault-classification and anomaly-detection models on acoustic and time-series sensor data from rotating machinery, reaching 93% fault prediction accuracy and cutting unplanned downtime by 67%.

Also: Research Intern at Nanyang Technological University (pre-trained MedBERT) & SSN College, where I built a 60,000-sentence Tamil emotion dataset and co-organized the Dravidian-CodeMix-FIRE 2021 shared task · Teaching Assistant at McMaster (DSP, Image Processing, Electromagnetics II) and the University of Moratuwa (Programming Fundamentals, Data Science, Deep Neural Networks).

Expertise

Engineering LeadershipTeam management (12 people incl. 2 tech leads & 2 PMs), hiring & performance management, technical mentorship, partnership engineering (IBM, NetApp, Intel, AMD, HP), budget ownership, product roadmap, cross-functional coordination.
AI/ML Systems & ArchitectureMulti-cloud (AWS, IBM Cloud) plus on-prem and edge inference; event-driven microservices; RAG, multi-agent orchestration (LangGraph), document intelligence, model quantization (GGUF, INT8/FP16), MCP integration.
Security & Multi-TenancyTenant isolation, RBAC, Active Directory/LDAP and SSO identity resolution (Okta, Entra ID), NFSv4 file ACLs propagated into vector retrieval, sandboxed execution of model-generated code (gVisor/Kata), air-gapped deployment.
Production Infrastructure & StackKubernetes, Kafka, FastAPI, OpenTelemetry, Prometheus/Grafana; Python, TypeScript, PyTorch, vLLM, OpenVINO; PostgreSQL, Redis, Milvus.

Selected Publications All on Scholar · 251 citations →

Google Scholar: h-index 7 · i10-index 6

  • SciRAG: A Retrieval-Focused Fine-Tuning Strategy for Scientific Documents
    Charangan Vasantharajan (Advisor: Kirubarajan Thia)
    Master's Thesis (MASc), McMaster University, 2025
  • MedBERT: A Pre-trained Language Model for Biomedical Named Entity Recognition
    C Vasantharajan, KZ Tun, H Thi-Nga, S Jain, T Rong, and CE Siong
    APSIPA ASC, 2022
    Code
  • Party Extraction from Legal Contract Using Contextualized Span Representations of Parties
    S Sivapiran, C Vasantharajan, and U Thayasivam
    RANLP, 2023
  • TamilEmo: Fine-grained Emotion Detection Dataset for Tamil
    C Vasantharajan, R Priyadharshini, PK Kumarasen, R Ponnusamy, and others
    Springer (SPELL), 2022
    Code
  • Adapting the Tesseract OCR Engine for Tamil & Sinhala Legacy Fonts and Creating a Parallel Corpus
    C Vasantharajan, L Tharmalingam, and U Thayasivam
    IALP, 2022
    Code
  • Towards Offensive Language Identification for Tamil Code-Mixed YouTube Comments and Posts
    C Vasantharajan and U Thayasivam
    SN Computer Science (Journal), 2021
    Code

Patents

System and method for Fine Tuning of Large Language Models

US Patent Application · 2025

B Sathianathan, L Jothivincent, S Tharumarasa, C Vasantharajan

Apparatuses, Systems, and Methods for Creating Artificial Intelligence Agents and Routes

US Patent Application · 2025

B Sathianathan, C Vasantharajan

Apparatuses, Systems, and Methods for Generating AI Workflows

US Patent Application · 2025

B Sathianathan, C Vasantharajan

Language Independent Textual Extraction

US App. 18/590,347 · 2024

B Sathianathan, C Vasantharajan, S Jacob

Open Source

Stepgate

MCP server that shows an agent one step at a time and moves on only when mechanical checks pass.

stars 9 · forks 1 · downloads 787

MedBERT

Pre-trained BERT for biomedical Named Entity Recognition.

HuggingFace · 569,801 downloads

NERP

Python package for fine-tuning transformer-based NER models.

PyPI · 71,746 downloads

Tamizhi-Net OCR

LSTM-enhanced Tesseract models for Tamil & Sinhala legacy fonts.

GitHub · IALP 2022

TamilEmo

Fine-grained emotion-detection dataset for Tamil.

CodaLab · Shared Task

Recognition

  • 2026CRN AI 100. Iterate.ai was named to the Top 20 Hottest AI Software Companies.
  • 2025Pinnacle AI Product of the Year, awarded to Generate.
  • 2025TMC Generative AI Product of the Year, awarded to Generate.
  • 2023McMaster Graduate Research Scholarship, a competitive grant supporting MASc research.
  • 2022Top 5 Finalist, HPCIC Innovation Challenge with team Medispeech@NTU.
  • 2018National merit scholarships: Suba Pathum, Mahapola, Ceylinco Pranama & Dialog awards for academic excellence.

Education

Master of Applied Science (MASc), Electrical & Computer Engineering

2023 to 2025

McMaster University, Hamilton, ON

  • Thesis: SciRAG — A Retrieval-Focused Fine-Tuning Strategy for Scientific Documents, advised by Prof. Thia Kirubarajan. RAG framework for scientific PDFs with custom LaTeX-equation parsing and domain-adapted LLM fine-tuning, reaching 70% GSM8k accuracy, 85% context recall, and 94% semantic similarity.

BSc in Engineering (Hons.), Computer Science & Engineering

2018 to 2023

University of Moratuwa, Sri Lanka · First Class Honours

  • Thesis: Entity Extraction System for Legal Contracts, on party extraction with contextualized span representations on a 1,000-contract annotated dataset, reaching 94.2% exact match (+6.2% over baseline).
© 2026 Charangan Vasantharajan