Moaz MuhammadAI Engineer

AI Engineer · Cairo · Arabic & English

معاذ محمد

I build AI that cites its sources or says it doesn't know.

LLM applications, RAG and document AI, shipped end to end in Python and TypeScript. Recent work: a live tender-compliance tool, an Arabic-first learning platform, and a clinical assistant for Egyptian-Arabic consultations.

Open to AI Engineer roles and contract workRemote · Cairo, UTC+2/+3
Can we mark “ISO 13485 certification” as met?
BidProof · grounded answer

Partial. The certificate is current, but its scope covers membranes for medical devices, not the respirators being offered.

tender.pdf · p.7iso-13485.pdf · scope
Partial · reviewer decides
Replay of real product behaviour1/3

Selected work

Products, shipped and measured.

Each one is live or public, and each one says plainly what it does not do yet.

01Document AI · Live productLive

BidProof

NTG Clarity · AI Product Internship · Sep 2026

Reads Arabic and English tenders and tells a bid manager whether the bid can actually be submitted. It refuses to call a requirement met unless an approved, cited, in-date document proves it.

  • Every requirement is a verbatim quote with its page; nothing reaches "satisfied" without approved evidence.
  • 1,309 of 1,528 production requirements quoted verbatim on the cited page.
  • Arabic and English tenders, PDF or DOCX, with selective OCR.
1,309 / 1,528
Quote audit
0.673 → 0.777
Extraction recall
BidProof submission checklist listing disqualifying tender requirements, each quoted with its source page
02EdTech platform · Live productLive

Eilm

Jul 2026 — present

An Arabic-first platform for independent academies. Each academy gets its own storefront, sells courses, teaches through lessons, assessments and live sessions, and takes payment into its own account.

  • Multi-tenant isolation enforced in the database with deny-by-default row-level security.
  • Each academy is merchant of record on its own Stripe or Paymob account; the platform never holds funds.
  • Lesson Q&A answers only from the transcript, with verbatim quotes, or refuses.
155
Database migrations
~840 + 91
Automated tests
Eilm English landing page — "Knowledge with your name. Progress with meaning."
03Clinical AI · Employment workIn rollout

AIDocFriend V2

Yashfii · Jun 2026 — present

A clinical assistant that turns an Egyptian-Arabic doctor–patient consultation into live differential diagnoses, editable SOAP notes, prescriptions, lab orders and medication-safety checks.

  • Streaming speech-to-text across four engines with automatic fallback.
  • LLM safety findings are filtered against the reconciled medication list before a doctor sees them.
  • Approved by the clinic owner as the platform going forward.
4
Speech-to-text engines
14
Architecture decisions
AIDocFriend V2 bilingual consultation workspace with live transcript, suggested questions and analysis panels
04AI agents · Open sourcePublic

PR Sentinel

Jun 2026

Multi-agent code review that runs in your own CI with your own LLM key, and shows which agent found what. Published on the GitHub Marketplace, with its real-world recall published too.

  • Four specialist analysts, three samples each, majority vote and a verifier before anything posts.
  • Findings whose quoted evidence isn't in the diff are dropped, mechanically.
  • 28% recall on 60 real bug-fix PRs, published beside the 91% seeded score.
~91% · 0 FP
Seeded benchmark
28% recall · 71% precision
Real bug-fix PRs
A PR Sentinel review comment listing critical, high, medium and low findings, each anchored to a file and line
05RAG platform · Open sourceLive

SecureAgentRAG

May — Jun 2026

Privacy-first English/Arabic agentic RAG where access control, citations and privacy are enforced by the system, not requested in a prompt.

  • Role-based access enforced inside the vector-database filter, for dense and sparse search alike.
  • Hybrid retrieval fused by reciprocal rank, then cross-encoder reranking.
  • Corrective re-retrieval that says "not found" rather than guessing.
~700
Automated tests
41
Architecture decisions
SecureAgentRAG public demo page describing RBAC at the vector layer, sensitivity routing, faithfulness gate and audit chain

Earlier work

How I build

Four habits that show up in every repo.

01

Evidence before generation

If a claim can't point to a source, the system says so. Refusal is a valid answer, not an error path.

BidProof's evidence lock · Eilm's quote-checked lesson Q&A

02

Publish the unflattering number

A real-world score you can improve beats a seeded score that flatters. Limitations ship with the result.

PR Sentinel: 28% real-PR recall shown beside 91% on fixtures

03

Tests before claims

Behaviour that matters is pinned by a test, including the model's refusals and prompt-injection cases.

~840 tests on Eilm · ~700 on SecureAgentRAG · 306 on BidProof · 256 on PR Sentinel

04

Decisions in writing

Every architecture call is recorded with its reason, so the next engineer (or agent) can disagree on purpose.

195 decision records on Eilm · 41 ADRs on SecureAgentRAG · 14 at Yashfii

For clients

What I can build for you.

Fixed scope, a small first milestone, and a measured result you can re-run, not just a demo.

Document Q&A with citations

RAG over your PDFs, policies, contracts or knowledge base. Every answer quotes its source; when the documents don't say, it says so.

First milestoneA prototype over 20–50 of your documents, with a 30-question accuracy report.

Document extraction, Arabic & English

PDF, DOCX and scanned files turned into structured data, each field carrying the quote and page it came from.

First milestoneExtraction for one document type, checked against a hand-labelled sample.

An LLM feature in your product

Added to your Python/FastAPI or TypeScript/Next.js codebase with cost caps, prompt-injection tests and a kill switch.

First milestoneOne feature behind a flag, tested, in your repository.

Evaluation for AI you already run

A labelled test set and a harness that tells you how often your assistant is right, before your users find out.

First milestoneA baseline precision/recall report and the three failure patterns that matter most.

Experience

Where the work happened.

Jun 2026 — Present
Cairo, Egypt · Hybrid

Yashfii

AI & Software Engineer (part-time)

Sole engineer on AIDocFriend V2, the successor to the team-built V1 clinical assistant.

  • Built a FastAPI/Python and Next.js/TypeScript monorepo that turns Egyptian-Arabic consultations into live differential diagnoses, editable SOAP notes, prescriptions and lab orders. The clinic owner approved V2 as the platform going forward.
  • Built a speech-to-text adapter over four engines (browser speech, Deepgram real-time streaming, Mistral Voxtral, local Qwen-based Egyptian-Arabic ASR) with automatic fallback, feeding staged live analysis.
Aug 2026 — Sep 2026
Cairo, Egypt · Hybrid

NTG Clarity

AI Product Engineer Intern — Team Lead

Led a three-person team in the AI Product Internship and wrote 196 of 197 commits for BidProof.

  • BidProof (live) reads Arabic and English tenders and marks a requirement "satisfied" only when approved, cited, in-date evidence proves it.
  • Raised extraction recall on 578 attorney-labelled CUAD clauses from 0.673 to 0.777 and parser verbatim accuracy to 0.975 (Arabic) / 0.979 (English); an audit of 1,528 production requirements found 1,309 quoted verbatim on the cited page and 7 not found.
Apr 2026 — Jun 2026
Remote

Product Manager Accelerator (PMA)

AI Engineer Intern

Owned the AI-engineering workstream on ComplyAI, a Claude-based AI-regulation compliance checker.

  • Raised rule-data QA from 70.3% to 99.8% (550/551 rules) across 11 regulatory rule sets and re-mapped 88 US rules mis-tagged with EU IDs.
  • Validated the assessment engine end to end against live Claude: 13/13 legal-reasoning checks and zero hallucinated rule IDs across 67 generated gaps.
Jan 2026 — Jun 2026
Remote

Lillup

AI Engineer Intern (part-time)

On-device multi-agent coaching assistant.

  • Built hybrid intent routing (LLM + rules), RAG with ChromaDB, BM25 and citations, output-safety gates, and pytest suites.
Jul 2025 — Aug 2025
Remote

Cellula Technologies

Machine Learning Engineer Intern

Classical ML under weekly mentor review.

  • Built taxi-fare regression and hotel-cancellation classification models (scikit-learn, XGBoost, Pandas).
Sep 2024 — Jul 2028
Expected 2028

Helwan University

BSc, Computer Science & Artificial Intelligence

Toolbox

What I reach for.

AI & retrieval

LLM applicationsRAGHybrid search & rerankingLLM evaluationLangGraph agentsDocument AI & OCRArabic NLPSpeech-to-text

Backend & data

PythonFastAPITypeScriptFastifyPostgreSQLSupabase (RLS)pgvectorQdrant

Product & delivery

Next.jsReactTailwindArabic RTL & i18nPlaywrightpytestGitHub ActionsDocker

Models

DeepSeekAnthropic ClaudeOpenAI-compatible APIsGroqOllama (local)Hugging Face

Contact

Have documents your AI can't be trusted with yet?

Send one example of a wrong answer, or the documents you want answered, and I'll tell you where I'd start. Open to AI Engineer roles, agency subcontracting and fixed-scope projects.

Email me