I'm an AI engineer who builds complete systems, not demos — I take an idea from raw data through model training and evaluation into a deployed service people actually use. I work across the whole modern AI stack and move fluidly between its layers: agentic LLM systems and RAG, deep learning for vision and language, speech, classical machine learning, and the backend and MLOps that put it all into production.
What I care about is building things that hold up in the real world. My agents run on genuine trust boundaries — the model proposes and the server verifies, so nothing reaches a user or a database ungrounded or unauthorized. I get ambitious systems to run cheaply and privately, down to full agentic pipelines on ≤ 4 GB VRAM with no cloud dependency at all. And I bring uncommon depth to Arabic-language AI, from dialect-aware speech to classical prosody — a space most tools treat as an afterthought.
Give me a problem that has to be accurate, auditable, and affordable at the same time — that's exactly where I do my best work.
My two largest bodies of work. The Local-First Research Assistant is fully open — code, design docs, and the evaluation harness that scores it. The voice gateway remains proprietary; its engineering is summarized here, and an architecture walkthrough is available on request.
|
A server-authoritative orchestration gateway for a real-time Arabic voice agent operating on live business data. The model may request work, but every read and write is validated, authorized, previewed, confirmed by voice, and committed by the server — never by the model.
|
Local-First Research Assistant — |
| Project | What it is | Highlights |
|---|---|---|
AI Music Generation Bot · private |
Suno-powered Telegram bot — generates songs, lyrics, covers, karaoke / stem splits, track extensions & music videos | ▶ Try on Telegram |
| Burn Detection & Severity Grading | SimSiam self-supervised pre-training → dual-head multi-task ResNet50 → FastAPI | Detection F1 0.9717 · Severity F1 0.8723 |
| Human Action Recognition | Temporal Shift Module (TSM) on UCF-101 with async FastAPI inference, JWT/RBAC, analytics | 92.4% Top-1 @ 27.6 FPS — 6.7× faster than I3D |
| Multimodal Deepfake Detector | Image + audio + video forensics with explainable evidence | AUC 0.917 / 0.989 / 0.925 |
| Arabic E-Commerce Trust Scorer | Live dual-fetch scraping + hybrid Regex/Gemini extraction → 100-point trust score | Real-time FastAPI microservice · Arabic |
| Arabic Poetry Analyzer | Triple-head AraBERT MTL (meter · rhyme · anomaly) + CATT diacritizer + local RAG | PCGrad · ~3.1K LOC microservice |
| Explainable Chest X-Ray Diagnosis | ImageNet vs RadImageNet interpretability study vs radiologist annotations | Grad-CAM++ · best BBox IoU 0.238 |
Hybrid retrieval (BGE-M3) · Reciprocal Rank Fusion · cross-encoder reranking · MMR · HyDE · Corrective RAG · sentence-level groundedness verification · semantic caching · multi-agent orchestration (planner · critic · executor) · sandboxed tool use · local LLMs (Qwen · Gemma · Phi-2)
ResNet · DenseNet · ConvNeXt · CoAtNet · U-Net · TSM · LSTM/BiLSTM · YOLOv8 · Multi-Task Learning · SimSiam (self-supervised) · supervised contrastive · PCGrad gradient surgery · focal / asymmetric / masked loss · mixed precision (AMP) · multi-GPU
Semantic segmentation · video understanding · Grad-CAM++ / EigenCAM / ScoreCAM · ELA & FFT image forensics
Arabic diacritization (CATT transformer) · dialect-aware intent routing · prosody & meter analysis · FinBERT-ESG domain tuning
WhatsApp +963 980 766 663 · Email yazanaboassa223@gmail.com · yazanaidev3@gmail.com
Building the next generation of intelligent systems — grounded, private, and production-ready.

