- Lead AI architecture for high-trust nuclear and legal workflows: production LLM/RAG over a 1M+ document archive, hybrid retrieval across Azure AI Search, Weaviate, FAISS, Qdrant, and Databricks, plus multi-agent systems with function calling and MCP servers.
- Own evaluation and reliability practice with MLflow, Ragas, LLM-as-a-judge, A/B and canary testing — measuring accuracy, groundedness, hallucination risk, latency, and cost before models reach operators and counsel.
- Fine-tune LLaMA-class models (including LoRA) for constrained on-site use; design guardrails and grounding for sensitive nuclear data in partnership with legal on FIPPA-aligned disclosure review.
- Guide data scientists and engineers on retrieval, evaluation, and pipeline delivery; ship product lines detailed in Projects (Sherlock, CoREY, ClearSight, TrendVision, Baywatch).
Ottawa, Canada
Momin Aziz
Senior Data Scientist / AI Engineer
8+ years designing, deploying, and evaluating production machine learning and LLM systems — from enterprise RAG over million-document archives to agentic workflows that teams can trust.
Production AI for high-trust environments
I build enterprise RAG and agentic AI systems over large document corpora — including a 1M+ document archive at Ontario Power Generation — using hybrid retrieval, multi-agent orchestration, function calling, tool-augmented reasoning, ReAct-style patterns, and MCP servers.
Reliability is part of the product: evaluation frameworks with MLflow, Ragas, and LLM-as-a-judge methods measure accuracy, groundedness, hallucination risk, latency, and cost, with A/B and canary testing for sensitive workflows.
I lead AI architecture and cross-functional delivery, optimizing models and pipelines for latency, cost, scalability, and dependable behavior where mistakes matter.
Selected roles
Roles across nuclear operations, real-time voice commerce, fintech content AI, supply-chain compliance, and privacy research — different domains, stacks, and delivery constraints.
- Created and maintained the lifecycle of an NLU engine that predicts items ordered by a customer in a drive-through, combining GPT-3, Hugging Face models, and a rule-based engine.
- Fine-tuned GPT-3 models and customized prompts for higher prediction accuracy; built a GPT-3–parallel prediction framework with open-source language models for on-premise deployment.
- Wrote GitLab CI/CD pipelines to test NLU performance and deploy across environments.
- Managed multiple annotation teams to produce high-quality labeled data with MTurk, LabelStudio, and DataSaur.
- Trained and improved audio-to-text predictions with Azure Automatic Speech Recognition (Azure Speech, AWS Polly); customized Azure Text-to-Speech for on-prem and cloud deployments.
- Managed AI-team data infrastructure on an event-driven backend with Kafka, Elasticsearch, AWS Redshift, and Docker.
- Developed Machine Learning APIs for text data on Google Cloud Platform with GCP Vertex.
- Built NLP mechanisms for content generation, chatbot replies, topic identification, and snippet extraction using GPT-3 and Hugging Face.
- Integrated ML APIs into Google add-ons, Google Docs, and Chrome extensions; wired services into an event-driven TypeScript backend.
- Used BigQuery as the GCP data warehouse for ML training pipelines; ran functional and load testing of ML endpoints with K6.
- Applied deep learning to extract critical information from image and text data; built systems that help humans assess compliance documents from arbitrary sources and languages.
- Deployed transformer pipelines on compliance documents — QA, classification, similarity, and NER — with BERT and Hugging Face; processed, cleansed, and verified data integrity for analysis.
- Served models on AWS SageMaker with serverless APIs (Lambda, Step Functions / State Machines, SNS, SQS) and deployed with CloudFormation and Python.
- Ran information-extraction R&D with OCR, spaCy, and Hugging Face; transformed and loaded text from PDF, images, and XML using AWS Textract, Translate, and Tesseract.
- Research Assistant at the University of Manitoba Data Security and Privacy Lab under Dr. Noman Mohammed: secure computation on healthcare data in centralized and federated settings, targeting privacy-preserving sharing of genomic and patient records.
- Applied Homomorphic Encryption, Garbled Circuits, Differential Privacy, and Intel SGX; designed secure protocols for string similarity queries and studied privacy/security of machine learning and textual healthcare data, including deep learning for EHR information retrieval.
- UC San Diego (Junior Specialist): GAN-based generation and classification of sensitive biomedical data; Homomorphic Encryption for multiparty private set intersection. EPFL (under Dr. Jean-Pierre Hubaux): secure logistic regression on genomic data with Homomorphic Encryption.
- iDASH Genomic Data S&P Competition runner-up — Task 1 (privacy-preserving dissemination) and Task 2 (secure collaboration), 2016.
Products and selected work
Personal products I am building, alongside high-impact systems from industry engagements.
A capture, search, and governance layer for AI interactions across coding assistants and providers. Sessions become searchable working context for individuals — and observability, cost attribution, DLP, and audit trails for teams that treat AI tools as critical infrastructure.
Native hooks for Claude Code, Gemini CLI, and Codex CLI record structured sessions with tool calls, diffs, and token usage; hybrid semantic and keyword search makes prior work reusable instead of lost inside vendor chat histories.
An AI-driven telephony and automation platform for small and medium businesses: smart voice assistants, AI chat, scheduling, Google Business Profile automation, and single-page website hosting in one place.
Built to reduce missed calls, cut manual booking and follow-ups, and strengthen local online presence for service businesses that cannot staff a full-time receptionist.
Document search engine for Ontario Power Generation spanning roughly one million documents that are periodically updated and critical to day-to-day nuclear operations. Documents are chunked and vectorized into a hybrid retrieval index that is heavily used internally for enterprise search and decision support.
The hard problems sit outside a one-time index build. Periodic corpus updates stress RAG pipelines — stale chunks, re-embedding cost, consistency between keyword and vector stores, and safe rollouts when the source of truth keeps changing. Extraction quality varies wildly across typewritten scans, modern office documents, dense tables, and technical drawings, so retrieval failures often start as OCR and layout failures rather than model failures.
Cost-estimation chatbot for retrieving, analyzing, and verifying project estimates on top of OPG’s enterprise document systems — with evaluation harnesses and guardrails suited to nuclear and legal workflows.
Improvements in retrieval and analytics were tied to estimated savings of $5–10M per year in production use.
Databricks-based AI document intelligence for Freedom of Information (FOI) workflows at Ontario Power Generation: SharePoint ingestion, PDF/OCR extraction, Azure OpenAI, structured metadata generation, and automated output delivery. Companion pipelines generate tables of contents for large unstructured packages by detecting document boundaries across emails, reports, tables, letters, presentations, and attachments.
The core capability is sensitive-data handling for disclosure review. The FOI redaction pipeline analyzes extracted text, table structure, and layout polygons together so exemption rules can be applied with layout awareness—not text alone—then produces structured redaction recommendations and writes annotations/redactions back to PDFs. Quality is measured against ground-truth samples with page-level precision/recall/F1, span-level ROUGE, confusion matrices, token and runtime tracking, and MLflow experiment logging to compare prompts and models.
Turns narrative operational records into structured trend signals, reviewable model outputs, and scheduled Databricks workflows. Focused on equipment, safety, operational, regulatory, and performance conditions in nuclear operations — combining data engineering (ingestion, text preparation, acronym expansion, scheduled pipelines) with AI engineering (embeddings, clustering, classification, and Azure OpenAI summaries) so records are easier to analyze and act on.
The core workflow family is SCR (Station Condition Record): prepare text from titles, condition descriptions, immediate actions, and recommended resolutions; embed and classify into business categories; cluster related records to surface emerging or recurring issues; and summarize clusters for review. Related pipelines cover ONC (Observation and Coaching), TPI (Trending and Performance Intervention), and VOT (Validation-of-Trend) summaries for human-performance and trend-validation records.
Backend and primary heuristic algorithm for nuclear waste-container movement optimization — minimum-move planning under stacking and age-based constraints to improve storage utilization. Recognized with a Gartner Innovation Award at Ontario Power Generation.
Background
-
Ph.D., Computer Science
University of Manitoba · Manitoba, Canada · 2017 – 2022
-
M.Sc., Computer Science
University of Manitoba · Manitoba, Canada · 2015 – 2017
-
B.Sc., Computer Science and Engineering
Bangladesh University of Engineering and Technology · Dhaka · 2009 – 2014
-
Gartner Innovation Award
Ontario Power Generation — Project Baywatch heuristic for nuclear waste container movement optimization
-
Amazon AWS research grant
2017 – 2022
-
Think Swiss Scholarship
2017
Selected papers
Top cited works in genomic privacy, secure computation, and healthcare NLP — full list on Google Scholar.
- SAFETY: Secure gwAs in Federated Environment Through a hYbrid solution
- Privacy-preserving techniques of genomic data — a survey
- De-identification of electronic health record using neural network
- PRESAGE: Privacy-preserving genetic testing via Software Guard Extension
- CPU and GPU accelerated fully homomorphic encryption
- Secure approximation of edit distance on genomic data
- Private and efficient query processing on outsourced genomic databases
- Differentially private medical texts generation using generative neural networks
- Aftermath of Bustamante attack on genomic beacon service
- Secure and efficient multiparty computation on genomic data
Let’s talk on LinkedIn
Prefer LinkedIn for introductions and collaboration — no forms, no inbox spam. GitHub and Google Scholar are available for code and publications.