AIDocFriend V2
A clinical assistant that turns an Egyptian-Arabic doctor–patient consultation into live differential diagnoses, editable SOAP notes, prescriptions, lab orders and medication-safety checks.

Browser speech, Deepgram real-time streaming, Mistral Voxtral, and a local Qwen-based Egyptian-Arabic model, behind one adapter.
Formal ADRs, plus implementation decisions recorded during the build.
The owner accepted V2 as the successor to the deployed, team-built V1. V2 deployment is waiting on production accounts.
API and web tests, strict type checks and real Supabase smoke runs (13 Jul 2026).
Problem
A doctor in Cairo consults in Egyptian Arabic, writes notes in English and has minutes per patient. The first version of the assistant proved the idea, but it kept tasks in memory, mixed doctors' data, and ran clinical rules as hard-coded heuristics that nobody wanted to maintain.
Approach
V2 is a polyglot Turborepo monorepo: a FastAPI/Python service for the AI work and a Next.js/TypeScript app for the doctor. Speech goes through a single speech-to-text adapter (browser speech first, Deepgram real-time streaming as the fallback, with Voxtral and a local Egyptian-Arabic Qwen model available). Analysis is staged, so the differential, SOAP note and prescription appear incrementally while the consultation continues.
Medication safety is an LLM pass for contraindications, precautions, monitoring and interactions, followed by a deterministic filter. Any finding for a drug that isn't on the reconciled active-medication list is dropped and logged. Doctor edits are re-checked against the prescription and saved as versions.
Under it: per-doctor authentication and encounter isolation, a leased PostgreSQL job queue with retries and stale-job recovery, and token, audio and cost telemetry with a per-encounter AI budget. Patient content never goes into operational logs.
Challenges
Real-time Arabic speech is unforgiving. Browsers time out, iOS needs its own audio path, and the streaming handshake has to wait for the server before audio flows. Live analysis needed throttling and stale-result rejection, so a slow answer never overwrites a newer transcript.
Outcome
I presented the architecture and its running costs to the clinic owner, who accepted V2 as the platform going forward. The predecessor V1 (built by a team) is deployed. V2's Docker and Cloud Run release path is scripted and waits on production accounts.
What I'd change
Replace "chosen by feel" with a proper word-error-rate benchmark across the speech engines on real (consented) Egyptian-Arabic audio, and verify the streaming microphone path on real Android and iOS devices before rollout.