Moaz MuhammadAI Engineer
← All work
AI agents · Open sourceLive

PR Sentinel

Multi-agent code review that runs in your own CI with your own LLM key, and shows which agent found what. Published on the GitHub Marketplace, with its real-world recall published too.

A PR Sentinel review comment listing critical, high, medium and low findings, each anchored to a file and line
~91% · 0 FP
Seeded benchmark

37 fixtures across 7 languages, with zero false positives on the clean-code controls.

28% recall · 71% precision
Real bug-fix PRs

60 merged bug-fix PRs from 20 repositories, reversed to reintroduce the bug, 3-run average. The honest hard number, not the flattering one.

~$0.004–0.01
Cost per review

Typical PR, depending on the provider route (DeepSeek V4 Flash to gpt-5-mini). Your key, your CI, no hosted service.

256
Automated tests

LLM and GitHub API fully mocked; no network in CI.

Problem

AI code reviewers tend to be noisy black boxes, and false positives are why teams uninstall them. Most also run on someone else's server with someone else's key.

Approach

A fan-out/fan-in LangGraph graph with no loops. Architecture, Security, Performance and Test analysts each run three samples in parallel and majority-vote. The merge step anchors every finding's quoted evidence to a real diff line, so a hallucinated finding is dropped by code, not discouraged by a prompt. A Verifier then confirms, rejects or downgrades each survivor against the code, and a Reviewer writes one prioritized comment.

It ships as a GitHub Action that speaks the OpenAI-compatible protocol with a configurable base URL. It runs in the user's own CI with their own key, and a dry_run mode estimates cost without calling a model.

Challenges

Two of them. First, prompt injection: a hostile PR must never steer the reviewer, so it runs on pull_request only (never pull_request_target), with minimal permissions and a structured-output boundary between diff text and instructions.

Second, honesty. Six proposed accuracy levers were measured against a multi-run real-PR benchmark. Most landed within run-to-run noise, and those ship off by default instead of being quietly switched on.

Outcome

Published on the GitHub Marketplace: ~91% with zero false positives on seeded fixtures, and 28% recall / 71% precision on 60 real merged bug-fix PRs, both published side by side. Backed by 256 tests.

What I'd change

The misses are context-dependent defects (a removed workaround, a teardown ordering bug) that a diff-only reviewer can't judge. The next lever is cross-file context, kept off until a multi-language, multi-run benchmark shows it actually helps.