BidProof
Reads Arabic and English tenders and tells a bid manager whether the bid can actually be submitted. It refuses to call a requirement met unless an approved, cited, in-date document proves it.

Production requirements quoted verbatim on the cited page (22 Sep 2026). 79 close matches, 133 from a DOCX with no page numbers, 7 not found.
578 attorney-marked clauses in the public CUAD corpus, after a prompt fault was traced by clause category.
Verbatim accuracy against real tenders, after column detection and Arabic ligature repair.
Support verdicts judged correct by hand. Small and not independent, so it is shown as a count, not a rate.
Problem
A public tender is fifty pages of prose with a few hundred obligations, and maybe sixty of them can end the bid on a technicality: sign every page, submit two envelopes, hold a certificate whose scope actually covers the goods. Missing one loses the contract, and nobody notices until the rejection letter.
AI summaries make this worse. They produce fluent prose with no citation, so the bid manager has to re-read the tender anyway.
Approach
Evidence before generation. The pipeline parses the tender (PDF or DOCX, Arabic or English, selective OCR) into pages that keep their source locator. It then extracts every requirement as a verbatim quote plus one testable obligation, classified by what failing it costs. Next it retrieves candidate evidence from the company's library (pgvector plus lexical search, approved and in-date documents only) and verifies each pair.
An evidence lock sits on top: nothing reaches "satisfied" without approved, cited proof, and a reviewer can't override it. A reviewer can only confirm, on the record, something no document can prove. A second model flags clauses that look like grounds for rejection when the extractor missed them, and readiness waits until a person settles each one.
Challenges
The failures weren't random, and that made them fixable. Early recall misses
clustered in specific clause categories, which pointed to a prompt fault rather
than a model limit. A genuine, current ISO certificate whose scope covered a
different product line was being accepted as proof. A scope rule now makes that
"partial", and on the affected tender readiness correctly flipped from ready to
blocked.
Outcome
Live on real tenders, with an invite-only workspace, roles, spend caps, an amendment diff, a tender calendar read from the document itself and compliance-matrix exports. I led a three-person team in NTG Clarity's AI Product Internship and wrote 196 of the 197 commits.
What I'd change
Verification precision rests on 13 hand-judged requirements, which is far too few for a procurement lead to trust as a rate. The next step is an independently labelled set of a few hundred requirement–evidence pairs, measured before any further prompt work.


