Machine Learning Ops2021 – 2022ML Engineer

ML Pipeline over a REST API

DVC, GitHub Actions, FastAPI, and Heroku CD

End-to-end ML pipeline with datasets and models on AWS S3, DVC for data versioning, GitHub Actions CI, FastAPI for inference, and Heroku for continuous deployment. Local and live API tests via Pytest.

The Problem & Engineering Constraint

The Core Challenge

Training artefacts, datasets, and inference code drifted independently without data versioning or a tested HTTP contract.
Technical Architecture & Approach

Engineering Solution & Implementation

Versioned data and models with DVC on S3, used GitHub Actions to create the virtual environment and run CI, exposed application functions over FastAPI, and deployed continuously to Heroku. Visualisation with Matplotlib and Seaborn.

Pipeline: fetch → preprocess → train / validate → inference artefact → registry
ML pipeline flowchart from raw data through training to model registry

View repository on GitHub

Measured Production Impact

Verified Outcomes & Deliverables

CI/CD path from commit to a live inference API.

Data and model versioning via DVC + S3.

Pytest coverage for both local and live API endpoints.

Technologies & Components

System Tooling & Technologies

AWS S3DVCGitHub ActionsFastAPIHerokuPytestPandas