Selected Projects
ChurnOps
Customer churn prediction and production-oriented MLOps.
ShippedProblem
Churn models only create value if they survive production — training must be tracked, predictions served reliably, and data drift caught before decisions silently degrade.
Engineering
- •Train Random Forest and Logistic Regression on Telco churn data (80/20 stratified split)
- •Track parameters and metrics per run in the churn-prediction MLflow experiment
- •Register versioned artifacts — model, preprocessor, reference distributions — synced from Hugging Face Hub on startup
- •Serve single and batch predictions through FastAPI: POST /predict, POST /predict_batch, GET /health, GET /metrics
- •Monitor PSI and Jensen-Shannon drift over a sliding window (default 500) of recent predictions from the SQLite log
Evidence
- •19-field Pydantic request validation with typed prediction, health, and metrics responses
- •Drift status thresholds (OK / WARNING / ALERT) with Slack webhook alerts on ALERT
- •Streamlit dashboard with Plotly gauge chart and live drift monitoring sidebar
- •Docker Compose deployment (API + MLflow server) with a persistent prediction-log volume
- •GitHub Actions CI: format check, pytest suite, and Docker build verification
- •Structured JSON logging to stdout
This project demonstrates production-oriented ML engineering: a trainable model, a containerized API, automated CI, and observability beyond accuracy metrics. Built from scratch with no boilerplate generators.
RAGNar
Document-based RAG / grounded question answering system.
ShippedProblem
Generic chatbots guess — teams need answers grounded in their own documents, so uploads must become cited, verifiable responses instead of plausible-sounding text.
Engineering
- •Ingest PDFs and text files with automatic chunking (POST /ingest)
- •Embed chunks with OpenAI text-embedding-3-small into a persistent ChromaDB store
- •Retrieve top-k chunks above a similarity threshold for each question (POST /ask)
- •Generate grounded answers with citations via gpt-4o-mini at temperature 0, returning answer + sources + grounded flag
- •Evaluate automatically with the eval/ suite for RAG quality
Evidence
- •Retrieval recall@5: 97.06% and grounding accuracy: 100.00% on the automated evaluation baseline
- •Answer quality of 4.88 with 0 unparseable responses in the same baseline
- •Methodology in the eval/ suite — reproducible with make eval; see the evaluation baseline in the repository README
- •Streamlit UI for document upload and Q&A over a FastAPI + ChromaDB backend
- •Docker Compose multi-service deployment with CI/CD via GitHub Actions
Full-stack RAG system that demonstrates production-ready retrieval-augmented generation — from document ingestion to grounded, cited answers. Built with automated evaluation, Docker deployment, and CI/CD.
Banks ETL
Installable data pipeline for bank data processing.
ShippedProblem
A coursework ETL only convinces when it ships like production software — installable, logged, tested, and containerized instead of a one-off script.
Engineering
- •Extract — scrape the archived Wikipedia table of largest banks by market cap
- •Transform — convert USD market cap to GBP/EUR/INR via a live rates API with CSV fallback
- •Load — write CSV and SQLite (Largest_banks table, replaced each run)
- •Query — three SQL queries answering London/GBP, Berlin/EUR, and New Delhi/INR rankings
Evidence
- •Installable CLI: banks-etl full run plus banks-etl-cli run / query --city subcommands
- •Live Frankfurter exchange-rate API with graceful fallback to a cached CSV
- •Structured logging to logs/code_log.txt and stdout
- •Hermetic pytest suite: 24 tests with fixtures and mocked HTTP, flake8 clean
- •GitHub Actions CI: lint, test, and coverage on push and pull request
- •Dockerized with named volumes persisting the database and logs
- •Modular typed package with single-source config in banks_etl/config.py
This project demonstrates classic data-engineering fundamentals — reliable extraction, reproducible transforms, and tested loads — packaged the way production Python ships: CLI, CI, containers, and docs.
Technical Skills
Production ML
Training, tracking, and monitoring models in production — from experiments to drift detection and observability.
LLM / RAG / AI Applications
Grounded language systems — retrieval, embeddings, and cited generation over your own documents.
Data Engineering
Extraction, transformation, loading, and querying — from web scraping to SQLite.
Engineering Infrastructure & Backend
The backend behind the models — APIs, containers, CI, testing, and dashboards.
Experience
Intern — Industrial Systems
ENIE (Entreprise Nationale des Industries Électroniques)|Sidi Bel Abbes, Algeria
March 2026 – April 2026
- •Collaborated with engineering teams to document industrial electronics testing workflows and identify process inefficiencies
- •Applied systematic troubleshooting to hardware-software integration issues, developing a methodical approach to debugging complex systems
- •Gained exposure to quality assurance protocols and cross-functional technical communication in a production environment
I bring the same testing, debugging, and documentation discipline to my ML work — validated code, tested pipelines, and tracked experiments.
Education
Bachelor's in Computer Science — AI Specialization
Sep 2022 – Jul 2027
Université Djillali Liabès, Algeria
Relevant coursework: Machine Learning, Deep Learning, Natural Language Processing, Computer Vision, Big Data Analytics, Linear Algebra, Probability & Statistics, Algorithms & Data Structures
Certifications
- Machine Learning SpecializationDeepLearning.AI2025
- Generative AI with Large Language ModelsDeepLearning.AIAugust 2026
- TensorFlow Developer Professional CertificateDeepLearning.AIAugust 2026
- Deep Learning SpecializationDeepLearning.AIAugust 2026
Also completed: The AI Engineer Path — Scrimba (July 2025) · The Data Science Course: Complete Bootcamp 2025 — Careers 365 (February 2025)
Let's build something useful.
I'm actively seeking AI/ML Engineering internships and junior AI/ML Engineer roles for 2026–2027. If you're working on NLP, Computer Vision, or applied LLM systems, I'd love to hear from you.
ismailferdi10042004@gmail.com
linkedin.com/in/ismail-ferdi-1b3a70290
GitHub
github.com/ismailferdi
Open to: Internships, junior roles, freelance ML projects, open-source collaborations, research assistant positions.
