From the Kernel to the Cluster

ACHRAFBELARIF

AI Infrastructure / Linux Engineer

>AI Infrastructure Engineering

@Oracle|Casablanca, Morocco

Scroll
01.

ABOUT_ME

I am an AI Infrastructure Engineer specializing in the reliability, qualification, and recovery lifecycle of bare-metal GPU/HPC clusters powering large-scale AI/LLM training capacity in a hyperscale data center, with hands-on Linux virtualization experience (QEMU/KVM/Libvirt) applied directly to qualifying AI infrastructure platforms ahead of rollout, and backend engineering experience building web services that integrate with that same hardware and systems infrastructure.

With a foundation in C/C++, Python, and Bash, I specialize in understanding complex system interactions across the AI infrastructure stack, from kernel internals and hypervisors down to RDMA/InfiniBand fabrics, NVLink topologies, and physical hardware, isolating genuine faults from transient automation issues and turning validation evidence into repair-ready technical records for AI data center capacity.

The views expressed on this website are my own and do not necessarily reflect the views of Oracle. Nothing here is written on behalf of my employer, and all technical detail is described in general terms only.

AI Infrastructure & GPU/HPC Reliability

Hardware Qualification & Triage: Validating bare-metal GPU/HPC hosts against automated validation tooling, isolating genuine hardware faults from transient automation/control-plane issues
RDMA & Fabric Diagnostics: InfiniBand/RoCE fault isolation, mapping logical failures to NIC, PCIe BDF, cable/transceiver, and top-of-rack switch-port paths
GPU & Fabric Health: NVLink/NVSwitch topology, Xid events, NVIDIA Fabric Manager, PCIe/AER kernel diagnostics
Hardware Recovery: Deep power cycle, reprovisioning, cable/NIC/GPU replacement, and RMA coordination with repair-ready evidence records
AI Data Center Design: Rack-scale GPU topology, liquid cooling fundamentals, firmware/new-capacity enablement, fleet lifecycle management

Backend & Systems Engineering

Backend Engineering: FastAPI, NestJS, Dropwizard, Oracle JET, and web services integrated directly with hardware and AI infrastructure systems
Databases: PostgreSQL, Oracle Database, SQL, Prisma
Systems Programming: C/C++, Java, JavaScript, low-level APIs, memory management, concurrency
Scripting & Automation: Python, Bash for AI infrastructure tasks and tooling
Container Technologies: Docker, Podman, container orchestration with Kubernetes
Infrastructure Tools: Git, Jira, CI/CD pipelines, monitoring systems

AI Infrastructure Virtualization

System Performance: Process scheduling, memory management, I/O subsystem tuning
Fabric & GPU Diagnostics: rdma link, mlxlink, ethtool, ibv_devinfo, lspci, nvidia-smi for AI fabric fault isolation
Firmware & Provisioning: PLDM firmware updates, NVMe provisioning, host discovery and lifecycle management for AI capacity
Virtualization Interfaces: KVM/QEMU, Libvirt, live migration — validated ahead of AI infrastructure GPU platform rollouts
Secure & Confidential Computing: Guest isolation, attestation workflows, confidential computing concepts for multi-tenant AI infrastructure
Automated Testing: Python/Bash/Avocado frameworks for AI infrastructure regression and release-readiness validation

Professional Focus

AI Data Center Reliability: Qualification and recovery lifecycle for hyperscale GPU capacity powering large-scale AI/LLM training
Systems Engineering: Designing robust, scalable AI infrastructure solutions
AI Platform Validation: Automated testing frameworks for AI infrastructure platforms, reproducible environments
Root-Cause Analysis: Debugging complex system failures across stack layers, from software down to physical fabric
Backend Web Development: Extended the Oracle Contributor Agreement (OCA) Signing Service (Dropwizard/Oracle JET) with new features, performance improvements, and bug fixes for external open-source contributor onboarding
Tool Development: Custom utilities for system monitoring, testing, and automation
Documentation & Knowledge Sharing: Technical writing, mentoring, best practices

Projects & Initiatives

Agentic AI for Validation Workflows: Agentic loop with custom tools and skills under a multi-agent architecture, applied to end-to-end virtualization validation (workplace project, details not public)
Linux Lab Infrastructure: Personal test environments for kernel modules, container networking
Automation Frameworks: Custom tools for environment setup, log analysis, metric collection
Systems Research: Deep dives into kernel features, performance optimization techniques
Open Source Contributions: Patches, documentation, community engagement

Professional Attributes

Systems Thinking: Holistic approach to complex technical problems
Deterministic Reliability: Focus on reproducibility, predictability in infrastructure
Rigorous Documentation: Clear, actionable technical communication
Trilingual Communication: Native Arabic, professional French/English fluency

Skills Matrix

AI Infrastructure · AI Infrastructure Virtualization · Backend Engineering

HPC & GPU Infrastructure

NVIDIA Data Center GPUsAMD Instinct AcceleratorsRDMAInfiniBand/RoCENVLinkNVSwitchPCIe

AI Data Center Design

AI/HPC Data Center ArchitectureRack-Scale GPU TopologyLiquid Cooling FundamentalsFleet Lifecycle Management

AI Infrastructure Virtualization

Linux/UnixOS InternalsQEMUKVMLibvirtLive MigrationFirmwareDockerPodmanKubernetes

Networking & Security

VLANRoutingFirewallVPNDNS/DHCPGuest IsolationAttestationConfidential Computing

AI & Automation

Agentic AIMulti-Agent ArchitectureAI Infrastructure Automation

Backend & Programming

CC++JavaJavaScriptTypeScriptPythonBashFastAPINestJSDropwizardOracle JETReact

Databases

PostgreSQLOracle DatabaseSQLPrisma

Testing, Diagnostics & Automation

Hardware Validation Test Suitesmlxlinkethtoolnvidia-smirdma linkibv_devinfolspciAvocadoGitJira

Languages

ArabicEnglishFrench
02.

PROJECTS

AI infrastructure automation, Linux virtualization tooling, and backend engineering. From an agentic AI framework driving virtualization validation to systems-level projects in C/C++.

01

Agentic Validation Framework

Multi-agent framework that plans and executes virtualization validation and testing tasks through an agentic loop with custom tools and skills.

02

Fleet Log & Debug Dashboard

Web dashboard for remote log and debug data collection across Linux fleet hosts, providing centralized log streaming, search, and diagnostic snapshots.

03

Job Application Tracker

Personal dashboard for organizing a job search: manual and feed-based role capture, AI-assisted drafting of application material, and a review queue where every message is approved by hand before it is sent.

04

ft_transcendence

Real-time multiplayer Pong platform with live chat, matchmaking, 42 API OAuth, 2FA, and spectator mode.

05

ft_containers

C++ STL containers (vector, map, stack, set) and iterators reimplemented from scratch with full standard compliance.

06

webserv

HTTP/1.1 server in C++98 with I/O multiplexing, CGI support, virtual hosts, and chunked encoding.

07

ft_services

Kubernetes cluster with Nginx, WordPress, Grafana, MySQL, InfluxDB, and FTPS orchestrated via k8s manifests.

08

ft_server

Debian Docker container running Nginx, WordPress, phpMyAdmin, and MariaDB from bare configuration.

09

minishell

POSIX-compatible Unix shell with pipes, redirections, environment variables, heredoc, and built-ins.