Personal project · WIPAsk about Brazilian tax law or live company data, and every claim points at the exact span or API response behind it. A second model audits each claim, deterministic gates audit the auditor, and a question the evidence cannot support gets refused. Built to practice RAG, tool calling and LLM evals.

eval at 13/13 across two model families · answers audited claim by claim
Personal project · WIPPull fields out of documents without reviewing every one by hand: a queue, retries guided by the model's previous error, and human review only where confidence drops. Built to practice dead-letter queues, versioned evals and LLM guardrails.

versioned eval — 100% field accuracy on a synthetic golden set
Personal projectReal-time 1v1 Pokémon battle simulator with a server-authoritative engine that validates every move. Built to practice WebSockets and a real deploy: AWS ECS Fargate through GitHub Actions.

engine covered by unit tests · deployed from CI