Ahmed Doghri — Senior AI Engineer
Senior AI Engineer with 10 years shipping production LLM and ML systems — RAG pipelines, multi-agent orchestration, LLM evaluation, and MLOps — in regulated domains like pharma, fintech, and B2B. 64 open-source projects with reproducible demos and benchmarks.
Résumé (PDF) · Classic one-page site · contact@ideaboundary.com
64 projects
- VectorMorph — Interpolate one SVG shape into another and export a real animated SVG. Arc-length resampling, area-ordered subpath matching, and rotational alignment take self-intersections from 83 to zero.
- ATS Proof Resume Generator — Scores whether an ATS can actually read your resume, then rewrites it. The audit is deterministic, runs offline with no API key, and benchmarks at 1.00 precision and recall against resumes with planted defects.
- E-book Converter: My Files, My Rules — A local e-book converter (EPUB, MOBI, PDF) that verifies the conversion actually worked. Compares chapters, word count, images, and metadata between source and output, because exit code 0 does not mean the table of contents survived.
- vllm-cost-router — You are paying flagship prices to answer "what time do you close?" This routes the easy stuff to the cheap model and keeps the big brain for what needs it. 73% cheaper, and it measures whether the routing is actually right: the first keyword router misrouted 100% of held-out hard requests.
- guardrail-gate — Your LLM leaks a phone number and invents a refund policy in the same sentence. This catches both, and is measured against the hallucinations that actually get through: the ones reusing the source's own words, where word overlap missed 6 of 8.
- tablextract — Your data is trapped in a PDF, wedged between three paragraphs of prose. This digs it out, cites the exact cell, and refuses when the answer isn't there. The first version answered "what is Grade1 for the moon?" with a citation.
- citebench — Your RAG cites a source and nobody checks if it says what you claimed. I measured how much reranking fixes it, got 62% to 88%, then renamed the documents and the entire gain vanished. It had been reading filenames.
- churnfm — The world changed and your churn model never got the memo. I measured 80% back to 89% with a drift detector, then found the detector never fires on pure concept drift; an outcome-based signal catches what it misses.
- taggate — An auto-tagger that guesses is wrong as often as it is overconfident. I measured 75% to 100% with confidence gating, then found it escalating 22 of 24 ordinary products because its keyword list was reverse-engineered from its own demo.
- agentmem — Most agent memory is a junk drawer that grows forever. This one has opinions, or was supposed to: its salience gate wrote 37% of pure filler to memory anyway, because "I" always looks like a proper noun. Fixed, with a frozen holdout to prove it.
- rubricagent — You wrote your eval rubric once and never checked if it works. I measured 0.77 to 1.00 learning the real one from outcomes, then found the grounding criterion correlated at exactly 0.000 on ordinary text. It was built to detect its own test data.
- clarifyrag — Most RAG agents confidently answer the wrong question. This one asks "which jaguar did you mean?" only when it actually needs to. 100% accuracy, fewer questions asked.
- speculabench — Everyone quotes a 2-3x speedup for speculative decoding and nobody shows the math. This models the accept and reject loop so you can find your sweet spot. 1.41x at 50% agreement, 2.86x at 90%.
- kvsqueeze — Long context has one villain: the KV cache that grows forever. At 40% of full cache, a heavy-hitter policy holds 57% recall while plain recency manages 15%.
- injectguard — Untrusted text tells your agent to ignore its instructions and it usually obeys. This catches it and shows exactly which rule fired. 100% precision and recall on a red-team corpus.
- toolrouter — Give an agent fifty tools and it starts guessing wrong with total confidence. This abstains on genuine ties instead. Catches 5/5 ambiguous queries at zero cost to accuracy.
- structstream — You asked for JSON, you got a code fence and a trailing comma. Raw parsing recovers 7% of realistic malformed output. One repair pass recovers 100%.
- chunklab — Everyone tunes the embedding model and forgets how they cut up the documents. Structure-aware chunking retrieves the whole answer 80% of the time vs 50% for naive fixed-size cuts.
- semanticentropy — A model that knows the answer says it five different ways. A model that's guessing contradicts itself. Consistent answers score 0.08 entropy, hallucinations score 0.90.
- contextpack — Long context costs you twice: to send it, and in latency to read it. Compress to 50% of the original and every load-bearing fact survives, with a real recall benchmark.
- agentbudget — Your agent doesn't crash when it fails, it loops forever inside its own budget. Step limits catch 0/3 stalled traces. Loop detection catches all 3.
- debatekit — One model gets one shot at its own blind spots. A panel of five that debates for two rounds goes from 57% solo accuracy to 80%, same underlying skill level.
- pendulumlab — A tiny robot control lab with real physics and a real reward curve. CEM tunes the controller from -54.30 return to 217.38, with no simulator install.
- motifdiff — A symbolic MIDI generator that grades its own output, and then a second pass that caught the grading itself cheating: half the "generated" melody was a hardcoded motif.
- connectpuct — A playable Connect Four game with a benchmarked PUCT agent. Its 10/10 record was against admittedly weak opponents; against real minimax search it's a genuine ~55% contest.
- chronopatch — A patch-based forecaster with conformal intervals. Found the 15% gain was measured on one series matching its own hardcoded assumptions; on a different series it drops to single digits.
- orthoshift — A causal benchmark where naive comparison is wrong by 1.812. Found the orthogonal DML method doesn't actually beat plain adjusted regression here, both just crush the naive baseline.
- fedcal — A non-IID federated learning benchmark. Found its client calibration step usually hurts the worst client rather than helping -- the published seed was a lucky draw out of 60.
- graphpulse — A graph anomaly detector that catches ordinary-degree nodes in the wrong neighborhood. Found its 0.938 AUC score was reading the ground-truth label; fair version scores ~0.66.
- proteinmask — A safe toy protein-like infilling benchmark. Found its own "random baseline" was a rigged formula guaranteed to score 0%; a real random guess scores ~5%.
- tabflowmini — A synthetic tabular data demo with metrics attached. Mean KS is 0.056, plan gap is 0.040, duplicate rate is 0.000.
- riskbandit — A contextual bandit with a real risk budget. Conformal selection drops violation rate from 0.733 to 0.007 while beating the safe baseline reward.
- cellcontext — The same knockout lands differently in different cells. Keep the cellular context and held-out response error drops 32.8%.
- foldcontact — A sequence can look protein-like and still ignore the fold. Turned out the 100% "satisfaction" score was a tautology; true-identity recovery is a real ~25% vs. ~8% naive.
- pangraphmap — The read is not broken. Your reference never contained its path. Turned out "25 of 25" was a tautology; with realistic sequencing error the real graph-vs-linear gain is ~58 points.
- methyloadapt — Six target-species labels are not a training set. The published "chance to 100%" was a lucky seed; fixed adaptation now beats target-only reliably across 90 seeds, never once worse.
- driftfilter — Deployment data moves while frozen classifiers stand still. Forward-only prototype filtering recovers 22.1 points of accuracy.
- taskrouter — Average three specialists and you erase their edge. Training-free routing removes 80.2% of static merge error on separated tasks.
- distractrack — Recency memory swaps identities when objects cross. Turned out "100%" was reading the ground-truth label directly; a genuine motion-only fix recovers a real, modest gain instead.
- d3video — Generated motion hides one derivative deeper. The published "46% to 100%" was close to best-case; a less tailored artifact still shows a real ~23-point gain.
- restem — One separation pass leaves music on the vocal. Turned out the 27.74 dB gain needed the interferer's exact frequency; a real frequency estimator now holds ~20 dB across dozens of frequencies.
- binauralbench — A clean stem can still collapse the room. Linked stereo separation preserves the source position and removes 99.8% of ILD error.
- spanjudge — The answer passed, but the agent doubled its latency and cost getting there. Replay the trace before release and turn production behavior into a CI gate.
- vrsbridge — Two VCF rows can describe one molecular change and still miss each other downstream. Normalize the allele and four records collapse into two computable variants.
- spatialniche — Coordinates become tissue context. Across 250 seeded permutations, the compact tumor neighborhood reaches a 5.10 z-score.
- structuregrade — A single confidence color hides weak loops. Parse every residue and the demo structure earns a C, with two residues below pLDDT 50.
- crisprradar — One guide, both strands, every NGG site. The 149-base demo finds one exact target and two ranked off-targets without sending a sequence anywhere.
- phenopacketlint — A packet can parse and still contradict itself. Semantic linting gives the demo a 100 quality score across three phenotype assertions.
- mcpinterlock — Agents should not inherit ambient authority. One policy check blocks the demo call for both an unsafe network target and an escaped output path.
- unlearnaudit — Forgetting without collapse is the test. Turned out "1.000 to 0.481" was a self-lookup artifact; the corrected audit finds no real leak either before or after retraining.
- videoprivacy — A missed frame should not reveal a face. Tracking fills the gap and redacts 10 regions across five frames with two stable identities.
- manifestlens — Open the credential, not just the pixels. The demo resolves one ingredient, three edit actions, a signer, and a valid hard binding.
- loudnessgate — Loudness is a delivery contract. The podcast demo measures -18.4 LUFS, identifies a 2.4 dB correction, and preserves peak headroom.
- audiocatalog — Exports multiply while recordings stay the same. Six pair checks find the one duplicate at 99.48% similarity across WAV and MP3 names.
- strandshift — One sequence, 18 equivalent views, a 0.3273 prediction range. The audit catches a model that learned the window instead of the biology.
- phenorank — The top disease scores 0.6334 with a 0.2554 margin and survives all three leave-one-phenotype-out trials.
- tensorwarden — Three real artifacts enter. SafeTensors passes; a dangerous pickle and a traversal archive are quarantined with three critical findings.
- cacheisolate — Shared caching reveals the secret prefix through an 86 ms timing gap. Selective isolation blocks the leak while preserving one safe cross-tenant hit.
- ragpoisonbench — Four poison documents collapse clean recall from 1.0 to 0.0. Provenance quarantine restores every answer without an LLM judge.
- splatgrade — Thirty-six Gaussians earn a D: two giants, three duplicate pairs, four low-opacity points, and two floaters are localized before rendering.
- physicsvideo — Five actual MP4 scenarios contain four planted physical faults. The benchmark finds all four with no false positive on the clean control.
- avsyncdoctor — Five flash/click pairs expose 120 ms of offset and 5,000 ppm of drift. FFmpeg repair brings remeasured skew below 40 ms.
- codecguard — A clean candidate passes four gates. A low-pass, stereo-collapsed build is blocked on spectrum, transients, width, and SNR.
- safepathshield — The nominal path spends 39 steps in collision. Seventy-five minimal interventions keep every shielded step safe and still reach the goal.