Work / Fig. 08 · LLM systems · Evaluation · Full-stack · 2026

DIV-1 — this portfolio, built as a production AI system

72/72 golden questions — routed correctly, plus 37/37 prompt-injection cases — enforced in CI on every push.

Fig. 08

DIV-1 — this portfolio, built as a production AI system

LLM systems · Evaluation · Full-stack · 2026

Problem

A portfolio that recruiters and engineers can question like a model has to be honest, cheap and dependable: it must never invent a claim, must keep working when models are rate-limited, and must not cost real money per visitor.

Approach

A deterministic router — keyword rules plus BM25 retrieval over a verified dossier — picks typed answer cards; a model only narrates over exactly those facts, raced across providers with a spend cap. Every answer records the facts it used, so a cached answer is invalidated only when one of its facts changes. Injection attempts are screened before any model; figures in model output are checked against the dossier. Built with evals in CI, an MCP server for AI screeners, and first-party analytics in Neon Postgres.

Results

  • 72/72 golden questions

    routed correctly, plus 37/37 prompt-injection cases — enforced in CI on every push

  • $0 for most answers

    presets, honest absences and cached repeats need no model; a narrated answer costs about $0.0002

  • Per-fact cache invalidation

    answers are invalidated only when a fact they cited changes — the EpiCache idea, running live

  • Queryable by AI agents

    an MCP server, llms.txt and a typed dossier.json, so screening assistants read verified data

How it works

  1. Question

    console · API · MCP

  2. Injection screen

    before any model

  3. Router

    rules + BM25 over the dossier

  4. Typed answer cards

    exact, always

  5. Narration if needed

    model race · number tripwire

  6. Recorded in Neon

    answer · trace · per-fact cache