For six years I have built enterprise AI for production, from multi-agent platforms to edge inference, and led the teams that ship it across cloud, on-prem, and air-gapped environments.
I am an Engineering Manager at Iterate.ai, where I lead a team of 12 across frontend, backend, DevOps, and AI/ML, including 2 tech leads and 2 project managers. We build Generate, the company's enterprise AI platform. I joined as a data scientist and moved into the manager role in under two years. Generate has since won AI Product of the Year from both Pinnacle and TMC, and it ships through partnerships with IBM, NetApp, Intel, AMD, and HP.
The work stays hands-on. I hold an MASc in Electrical & Computer Engineering from McMaster University, where my thesis was SciRAG, a retrieval-focused fine-tuning strategy for scientific documents. I have 4 filed patents and peer-reviewed publications with 251 citations.
RAG · Multi-Agent Systems · Agent Runtimes & Memory · Edge & On-Prem Inference · Enterprise AI Security · Team Leadership
Generate is Iterate.ai's enterprise AI platform. It runs agentic workflows and document intelligence over a company's private data. I have owned its architecture since it was a 5-engineer edge project, and I own it now as a multi-tenant SaaS product that also ships air-gapped into regulated industries. It is the same system in both places, which is most of what makes it hard.
The architecture
Generate runs as event-driven microservices that deploy identically to public cloud, customer-managed on-prem, and edge hardware, and air-gapped installs get the same feature set with no outbound network. That constraint shaped nearly every decision. Model serving is hardware-agnostic across the major accelerator vendors. Low-bit quantization keeps capable models running on consumer-grade machines, so private document search can run entirely on a user's own laptop.
An extraction pipeline handles the documents that defeat naive retrieval: nested tables, embedded visuals, and mixed formats. Accuracy is 94%, and a patent-filed table-extraction method reaches 96% precision on tables. Agent sessions run 30 to 40 minutes and outgrow even a million-token context window. A graph-based memory layer carries user context across them by semantic recall, compacts the working context as it fills, and narrows oversized tool output recursively instead of truncating it.
Agents write and run code, so every session executes inside its own kernel-isolated sandbox and a compromised session cannot reach another tenant's. Enterprise identity resolves through to file-level permissions, which means retrieval is filtered by what each user is cleared to open rather than trimmed after the fact. Administrators bind a model to each AI service, and every binding passes a standardized compatibility suite before it can be used. The platform holds 99.8% uptime at 50,000+ requests/second. An architecture consolidation with connection pooling cut infrastructure cost 33% and improved performance 4×.
What it produced
DealsTwo enterprise contracts. I delivered on the engineering commitments behind both.
Reach10,000+ Intel AI PC unitsshipped via the Intel Software Advantage Program, and demoed at the Intel Vision 2024 keynote.
PartnerIBM, NetApp, Intel, AMD, HP. The healthcare launch with Hutchinson Regional Healthcare as first customer was featured in IBM's product blog.
StorageNetApp ONTAP integrationcovering snapshot and SnapMirror through consistency groups, ACL-aware access to bound NFS shares, and regional indexing that keeps customer data in-region.
IP3 patents filed from the platform work, covering agent creation and routing, AI workflow generation, and LLM fine-tuning.
Scaled the team from 6 to 12 across frontend, backend, DevOps, and AI/ML, including 2 tech leads and 2 project managers. I led 6 hires and promoted 2 ICs to Senior Applied AI Engineer. I also mentored 4 engineers and 3 interns, one of whom now owns the agents platform.
Led the engineering behind two enterprise contracts for Generate. The product went on to win AI Product of the Year from both Pinnacle and TMC.
Own engineering for Generate across cloud, on-prem, and edge (AWS, IBM Cloud, Intel/AMD/NVIDIA), so it ships both as SaaS and as an air-gapped enterprise install.
Architected an 8-microservice event-driven platform at 99.8% uptime and 50,000+ requests/second at peak, with multi-terabyte indexing.
Built a document intelligence pipeline for tables, visuals, and mixed formats. Optimizing the SQL and vector database integration got it to 94% extraction accuracy and 400% faster responses.
Own the cloud and tooling budget: cut infrastructure cost 33% and improved system performance 4× through architectural consolidation, PgBouncer connection pooling, and edge-inference optimization.
Launched Generate for healthcare revenue recovery via an IBM reseller partnership with Hutchinson Regional Healthcare, featured in IBM's product blog.
Engineering Manager (Contract)
Sep 2024 to May 2025
Iterate.ai, Remote, Canada
Took ownership of a 5-engineer team and defined the roadmap that made Generate Iterate.ai's primary enterprise revenue product.
Built an event-driven microservices platform (Kafka, FastAPI, LangGraph) for high-throughput AI workloads, and set up the engineering operating model the full-time team later inherited.
Applied AI & Research
Data Scientist
Nov 2023 to Sep 2024
Iterate.ai, Remote, Canada
Led a 5-engineer team building edge-deployed AI for private document search on Intel AI PCs. It shipped to 10,000+ units through the Intel Software Advantage Program.
Built an Outlook RAG plugin for email automation (85% time savings, 92% user satisfaction) and presented it at the Intel Vision 2024 keynote.
Built a novel table-extraction system (Non-Maximum Suppression + Vision Language Models) at 96% precision; filed US Patent App. 18/590,347.
Fine-tuned an LLM with 4/8-bit quantization for a food-ordering chatbot: order accuracy up 40%, unintended responses down 78%.
Earlier ML Engineering
2021 to 2023
ExentAI & Simplyfai, Sri Lanka (Remote)
ML Research Engineer, ExentAI: built a newspaper digitization service processing 65,000+ historical documents with multilingual OCR (Tamil/Sinhala/English) on AWS with Docker & Kubernetes. It handled 10TB+ of data at 45% lower memory usage. Also shipped LLM pipelines for sentiment, emotion, and hate-speech detection.
Engineering Intern, Simplyfai: built fault-classification and anomaly-detection models on acoustic and time-series sensor data from rotating machinery, reaching 93% fault prediction accuracy and cutting unplanned downtime by 67%.
Also: Research Intern at Nanyang Technological University (pre-trained MedBERT) & SSN College, where I built a 60,000-sentence Tamil emotion dataset and co-organized the Dravidian-CodeMix-FIRE 2021 shared task · Teaching Assistant at McMaster (DSP, Image Processing, Electromagnetics II) and the University of Moratuwa (Programming Fundamentals, Data Science, Deep Neural Networks).
@article{Vasantharajan2022MedBERTAP,
title={MedBERT: A Pre-trained Language Model for Biomedical Named Entity Recognition},
author={Charangan Vasantharajan and Kyaw Zin Tun and Ho Thi-Nga and Sparsh Jain and Tong Rong and Chng Eng Siong},
journal={2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)},
year={2022},
pages={1482-1488}
}
@inproceedings{sivapiran-etal-2023-party,
title = {Party Extraction from Legal Contract Using Contextualized Span Representations of Parties},
author = {S Sivapiran, C Vasantharajan, and U Thayasivam},
booktitle = {Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing},
month = {sep},
year = {2023},
address = {Varna, Bulgaria},
publisher = {INCOMA Ltd., Shoumen, Bulgaria},
pages = {1085--1094},
}
@InProceedings{10.1007/978-3-031-33231-9_3,
author={Vasantharajan, Charangan and Priyadharshini, Ruba and Kumarasen, Prasanna Kumar and others},
title={TamilEmo: Fine-grained Emotion Detection Dataset for Tamil},
booktitle={Speech and Language Technologies for Low-Resource Languages},
year={2023},
publisher={Springer International Publishing},
pages={35--50}
}
@article{Vasantharajan2021AdaptingTT,
title={Adapting the Tesseract Open-Source OCR Engine for Tamil and Sinhala Legacy Fonts and Creating a Parallel Corpus for Tamil-Sinhala-English},
author={Charangan Vasantharajan and Laksika Tharmalingam and Uthayasanker Thayasivam},
journal={2022 International Conference on Asian Language Processing (IALP)},
year={2021},
pages={143-149}
}
@article{Vasantharajan2021TowardsOL,
title={Towards Offensive Language Identification for Tamil Code-Mixed YouTube Comments and Posts},
author={Charangan Vasantharajan and Uthayasanker Thayasivam},
journal={SN Computer Science},
year={2021},
volume={3}
}
Patents
System and method for Fine Tuning of Large Language Models
US Patent Application · 2025
B Sathianathan, L Jothivincent, S Tharumarasa, C Vasantharajan
Apparatuses, Systems, and Methods for Creating Artificial Intelligence Agents and Routes
US Patent Application · 2025
B Sathianathan, C Vasantharajan
Apparatuses, Systems, and Methods for Generating AI Workflows
Master of Applied Science (MASc), Electrical & Computer Engineering
2023 to 2025
McMaster University, Hamilton, ON
Thesis: SciRAG — A Retrieval-Focused Fine-Tuning Strategy for Scientific Documents, advised by Prof. Thia Kirubarajan. RAG framework for scientific PDFs with custom LaTeX-equation parsing and domain-adapted LLM fine-tuning, reaching 70% GSM8k accuracy, 85% context recall, and 94% semantic similarity.
BSc in Engineering (Hons.), Computer Science & Engineering
2018 to 2023
University of Moratuwa, Sri Lanka · First Class Honours
Thesis: Entity Extraction System for Legal Contracts, on party extraction with contextualized span representations on a 1,000-contract annotated dataset, reaching 94.2% exact match (+6.2% over baseline).