Skip to content
View aliduabubakari's full-sized avatar
💭
I may be slow to respond.
💭
I may be slow to respond.

Block or report aliduabubakari

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
aliduabubakari/README.md
Alidu Abubakari — Applied AI and Semantic Data Engineer

LinkedIn Hugging Face ORCID Medium Email Profile views

Hello — I build AI and data systems that can be trusted

I am a Computer Science PhD candidate at the University of Milano-Bicocca. My work brings together semantic data engineering, knowledge graphs, agentic AI, workflow automation, research software, and systematic evaluation.

From intent to evidence: natural-language requirements → validated specifications → executable workflows → evaluation and recovery

I develop Python systems that transform heterogeneous data and natural-language requirements into structured, validated, and reusable software and data products.

🧠 Applied AI 🕸️ Semantic data ⚙️ Reliable workflows
Agentic systems, structured generation, tool use Knowledge graphs, entity linking, semantic enrichment Specifications, code generation, evaluation, recovery

What I am focused on

  • Semantic data and knowledge graphs: entity linking, schema alignment, metadata enrichment, RDF, OWL, SPARQL, Neo4j, Cypher, and semantic search.
  • Applied and agentic AI: LangGraph, LLM APIs, structured generation, multi-agent orchestration, tool use, validation, and bounded repair.
  • Workflow and data engineering: interoperable pipeline specifications, code generation, FastAPI services, PostgreSQL, Kafka, Spark, Airflow, and containerized execution.
  • AI evaluation and reliability: execution-based benchmarks, contract validation, failure analysis, confidence handling, provenance, reproducibility, and human review.

Featured systems

Project What it demonstrates
inLUMEN Code Generation Service Authenticated FastAPI service that generates and validates executable pipeline components, including Python code, dependencies, Dockerfiles, runtime manifests, and validation reports.
Entity Linking Agent LangGraph-based multi-agent system for contextual entity linking across multiple knowledge sources, with confidence scoring, execution monitoring, and asynchronous APIs.
SemTUI Semantic enrichment platform for transforming heterogeneous tables into structured, knowledge-graph-aligned data products.
Notebook Provenance Static analysis of computational notebooks using AST parsing, data-flow graphs, semantic classification, explainability, and column-level lineage.
PipeSpec + OrchSpec Structured, orchestrator-independent pipeline specifications with deterministic validation and compilation to multiple workflow platforms.
Real-Time Fraud Detection Forked reference implementation of a streaming ML pipeline using Kafka, Spark Streaming, Airflow, MLflow, MinIO, PostgreSQL, Redis, and Docker Compose.

Current engineering themes

Represent  ──►  Generate  ──►  Validate  ──►  Execute  ──►  Evaluate  ──►  Recover
   data          code          contracts      workflows      evidence       safely
  • Reliable workflow generation: moving beyond syntactically plausible code toward contract-aware generation, deterministic validation, execution, and governed recovery.
  • Semantic enrichment: connecting heterogeneous scientific and urban data to shared concepts, identifiers, metadata, and graph structures.
  • Provenance and reproducibility: extracting lineage from computational notebooks and producing auditable data and software artifacts.
  • Sustainable AI-assisted pipelines: evaluating generated workflows for runtime behaviour, quality, and energy and carbon impact through PISCES.

Technology palette

Python FastAPI LangGraph Neo4j PostgreSQL Elasticsearch Docker Apache Kafka Apache Spark Apache Airflow GitHub Actions

Languages: Python, SQL, Bash
Applied AI and ML: LangGraph, LLM APIs, Hugging Face, scikit-learn, embeddings, semantic search
Data and graph systems: PostgreSQL, PostGIS, Elasticsearch, Neo4j, RDF, OWL, SPARQL, Wikidata
Data and software engineering: FastAPI, Flask, Docker, Kafka, Spark, MLflow, MinIO, Redis
Workflow orchestration: Airflow, Prefect, Dagster, Argo Workflows
Interfaces: REST APIs, OpenAPI, Streamlit, Gradio

Selected machine-learning and analytics projects
  • Sepsis Classification with FastAPI — Imbalanced classification, model comparison, hyperparameter tuning, F1/AUC evaluation, and deployment through FastAPI, Streamlit, Docker, and Hugging Face Spaces.
  • Expresso Telco Churn Classification — Customer-churn modelling with data preparation, feature engineering, classification, model evaluation, and business-oriented retention insights.
  • Indian Startup Data Analysis — Exploratory analysis, statistical hypothesis testing, visualization, and interpretation of investment patterns across industries, locations, and funding stages.

Domain experience

My work spans scientific data infrastructures, smart grids, digital substations, urban and environmental data, health data, entity resolution, and AI-assisted workflow engineering.

Before my PhD, I worked on IEC 61850 digital-substation interoperability at KEPCO Research Institute and smart-metering infrastructure at Landis+Gyr.

Open-source engineering and research

I build reusable software, documented APIs, evaluation frameworks, technical demonstrations, and reproducible experiments through international research and engineering collaborations.

Selected publications and research outputs are available through my ORCID.

Technical writing

I write about applied machine learning, semantic data systems, AI-assisted workflows, and practical deployment on Medium.

Let us connect

I am open to applied AI, semantic data, knowledge-graph, data-engineering, AI evaluation, and technical solutions roles in Europe.

Connect on LinkedIn Send an email Build with structure, validate with evidence

Popular repositories Loading

  1. Sepsis-Classification-with-FastAPI Sepsis-Classification-with-FastAPI Public

    This project is focused on the accurate and efficient classification of sepsis cases using the FastAPI framework. Sepsis is a critical medical condition that requires prompt identification and trea…

    Jupyter Notebook 14 4

  2. churn-prediction-with-gradio churn-prediction-with-gradio Public

    This repository contains code and resources for building a churn prediction model using machine learning techniques, and deploying it with Gradio for a user-friendly interface. Gradio is used to cr…

    Jupyter Notebook 2

  3. Streamlit-grocery-sales-prediction-app Streamlit-grocery-sales-prediction-app Public

    Building a machine learning (regression model) webapp for grocery sales prediction using streamlit

    Jupyter Notebook 2 1

  4. Azubian_forecasting_Prediction Azubian_forecasting_Prediction Public

    The objective of this challenge is to create a model to forecast the number of products purchased per week per store over the next eight weeks, for grocery stores located in different areas in the …

    Python 2

  5. California-Housing-Price California-Housing-Price Public

    The purpose of this project is to build a machine learning model of housing prices in California using the California census data. This data has features such as population, median income, median h…

    Jupyter Notebook 1

  6. Music-Recommender Music-Recommender Public

    An unsupervised learning model which analyses music genre and gives predicts genre based on Decision Tree Classifier

    Jupyter Notebook 1