Moaz Muhammad
AI Engineer — LLM Applications, RAG & Document AI
Summary
AI engineer who builds LLM products that cite their sources or refuse to answer. Sole or lead engineer on live Arabic/English products (tender compliance, online academies), a clinical-documentation assistant and an open-source code-review agent, in Python and TypeScript on PostgreSQL, with tests and published evaluations, unflattering numbers included.
Experience
Yashfii
Jun 2026 — PresentAI & Software Engineer (part-time) · Cairo, Egypt · Hybrid
- Built a FastAPI/Python and Next.js/TypeScript monorepo that turns Egyptian-Arabic consultations into live differential diagnoses, editable SOAP notes, prescriptions and lab orders. The clinic owner approved V2 as the platform going forward.
- Built a speech-to-text adapter over four engines (browser speech, Deepgram real-time streaming, Mistral Voxtral, local Qwen-based Egyptian-Arabic ASR) with automatic fallback, feeding staged live analysis.
- After the LLM medication-safety pass, added a deterministic filter that drops any finding for a drug not on the reconciled medication list; plus per-doctor isolation, a leased PostgreSQL job queue, versioned clinical edits and per-encounter AI cost budgets; 14 ADRs.
NTG Clarity
Aug 2026 — Sep 2026AI Product Engineer Intern — Team Lead · Cairo, Egypt · Hybrid
- BidProof (live) reads Arabic and English tenders and marks a requirement "satisfied" only when approved, cited, in-date evidence proves it.
- Raised extraction recall on 578 attorney-labelled CUAD clauses from 0.673 to 0.777 and parser verbatim accuracy to 0.975 (Arabic) / 0.979 (English); an audit of 1,528 production requirements found 1,309 quoted verbatim on the cited page and 7 not found.
- Pipeline: PDF/DOCX + OCR, extraction, pgvector + lexical retrieval, verification, evidence lock; Fastify, React, Supabase PostgreSQL, Cloudflare R2, QStash; 306 automated tests.
Product Manager Accelerator (PMA)
Apr 2026 — Jun 2026AI Engineer Intern · Remote
- Raised rule-data QA from 70.3% to 99.8% (550/551 rules) across 11 regulatory rule sets and re-mapped 88 US rules mis-tagged with EU IDs.
- Validated the assessment engine end to end against live Claude: 13/13 legal-reasoning checks and zero hallucinated rule IDs across 67 generated gaps.
Lillup
Jan 2026 — Jun 2026AI Engineer Intern (part-time) · Remote
- Built hybrid intent routing (LLM + rules), RAG with ChromaDB, BM25 and citations, output-safety gates, and pytest suites.
Cellula Technologies
Jul 2025 — Aug 2025Machine Learning Engineer Intern · Remote
- Built taxi-fare regression and hotel-cancellation classification models (scikit-learn, XGBoost, Pandas).
Projects
Founder and sole engineer of a live, Arabic-first, multi-tenant learning platform (749 commits): 155 migrations, deny-by-default row-level security, ~840 automated tests plus 91 database test files, academy-owned Stripe/Paymob payments, and lesson Q&A that answers only from the transcript or refuses.
Multi-agent code review on the GitHub Marketplace (4 analysts × 3 self-consistency samples, evidence anchoring, verifier), ~$0.01/review with the user's own key; 256 tests. 28% recall / 71% precision on 60 real bug-fix PRs, published next to 91% on seeded fixtures.
Privacy-first English/Arabic agentic RAG: nine-node LangGraph pipeline, hybrid dense + sparse retrieval with RRF, cross-encoder reranking, RBAC inside the vector filter, enforced citations; ~700 tests, 41 ADRs. Fine-tuned BGE-Reranker-v2-M3 (NDCG@10 0.774 → 0.790), shipped opt-in.
Skills
AI / LLM
Engineering
Quality & delivery
Education & languages
Helwan University — BSc, Computer Science & Artificial Intelligence
Sep 2024 — Jul 2028 · Expected 2028Arabic (Native) · English (Professional working proficiency)