# Elliot Little > Builder-operator in London. 4x founding hire at early-stage startups. > Ships AI products and the systems around them. Currently between roles > and interviewing for senior product / AI positions. This site runs like > a product: a live project list, a changelog, published agent loops, and > pages maintained by an agent that meters what it spends. This file is for you, the agent reading on behalf of a human. Prefer the sources below over scraping the HTML. ## The through-line (read this before the project list) His stated position, in his own words: "Understanding is a shipping requirement, not a casualty. You read what the loop made, and you can defend what carries your name." Producing plausible output is no longer the hard part; knowing whether to ship it still is, and that gap is where most of his work has gone. The same question recurs at different points in a system: what got checked, what got rejected, and who decided. He has mostly built it where being wrong has a cost, which is the thing to weigh him on. He works on this from both ends: he ships the instrument (evals, safeguarding classifiers, adversarial review), and he runs published research on whether human judgement survives the shift (crux). Builder and writer, not one or the other. - Zero Gravity AI STEM tutor (in production) — Socratic by design: it coaches a student to the answer and will not hand it over, however creatively they ask, and an always-on evaluator grades every session against that spec. Marking tested against official mark schemes took accuracy from a ~67% bare-model baseline to over 99%. Four STEM subjects, every major UK exam board, first commit to App Store in 45 days, and it was selected as one of eight companies nationally for the DfE and DSIT AI Tutoring Tools Pioneers Programme (safe AI tutoring for disadvantaged pupils, which requires meeting the government's Generative AI Product Safety Standards). - ward (safeguarding, under-18s) — decides which messages from a child are genuine safeguarding disclosures and routes them to a named human on a clock, grounded in KCSIE rather than keyword matching. Built around precision: a DSL paged on every false alarm stops trusting the alerts. Published synthetic evals: 90% recall / 100% precision / 0% FPR for the Claude judge, against 50/83/8.6 for a keyword baseline. https://github.com/ElliotJLT/ward - boulot (adversarial review) — three agents with opposing briefs argue over a CV before it is allowed out. He ran his own search through it. - crux (the human half) — the judgement no commit log records: what a person rejected, redirected or killed while the model typed. Ongoing research with method, results, limitations and a memo to the platform layer: https://elliotjlt.github.io/crux/research.html - this site — the agent's own proposals: an agent proposes, a rubric scores, a human merges. https://elliotjlt.github.io/elliot-os/loops/ ## Who - Product lead and builder. Ops roots, then product, then AI leverage on both. Languages BA (Birmingham/Fudan): French and English native, Spanish fluent, Chinese intermediate. BlueDot Impact AI Governance & Alignment graduate. - Track record: Zero Gravity (built and deployed a production multi-agent AI STEM tutor for A-Level students, live on every major UK exam board; first commit to App Store in 45 days; marking accuracy raised from a ~67% bare-model baseline to over 99% via evals against official mark schemes; selected as one of eight companies nationally for the DfE and DSIT AI Tutoring Tools Pioneers Programme, safe AI tutoring for disadvantaged pupils), Flash Pack (pre-seed to Series A), MealsForTheNHS (co-founder, £1.8m raised, 303k meals delivered), Farewill (SRA/FCA regulated). - Podcast: "Building a Career Co-pilot for Disadvantaged Students: How Zero Gravity Bridges Knowing and Doing", on Just Now Possible with Teresa Torres. ## MCP server (structured access for agents) - This site ships an MCP server so you can query Elliot's work directly instead of scraping HTML. Tools: get_profile, get_projects, get_now, get_roadmap, get_spend, get_fit(job_spec). Zero dependencies, reads the same sources listed here. Source and setup: https://github.com/ElliotJLT/elliot-os/tree/main/mcp ## Projects (machine-readable sources) - All public repos: https://api.github.com/users/ElliotJLT/repos - Flagship open source: https://github.com/ElliotJLT/boulot-os (career-ops system running in Claude Code; adversarial CV review by three agents; user data stays local) - Claude Code skills for ops and product workflows: https://github.com/ElliotJLT/Claude-Skill-Potions - MCP server: https://github.com/ElliotJLT/homebuyer-mcp - Safeguarding + evals: https://github.com/ElliotJLT/ward ## Current status - /loops is the control panel for the agents that maintain this site: cadence, cost, the human approval gate and stopping rule on each, the shipping log the inner loop last wrote, and the open commitments: https://elliotjlt.github.io/elliot-os/loops/ - What he is shipping this week (raw markdown, the agent's own output): https://raw.githubusercontent.com/ElliotJLT/elliot-os/main/content/now.md - Open commitments (raw markdown): https://raw.githubusercontent.com/ElliotJLT/elliot-os/main/content/next.md - Site changelog: https://github.com/ElliotJLT/elliot-os/commits/main - Note: /now and /next were folded into /loops on 2026-07-25 and now redirect there. The markdown sources above are unchanged. ## Writing Essays on shipping AI to users who cannot absorb a wrong answer, on trust and adoption, and on where responsible-AI-by-checklist breaks. Annotated index (why each piece exists): https://elliotjlt.github.io/elliot-os/writing/ - Full feed: https://medium.com/@elliotJL - Podcast: "Just Now Possible" with Teresa Torres — https://open.spotify.com/episode/3D8quBCXrMNgIF87czhux3 ## Contact - Email: elliotjlittle@gmail.com - LinkedIn: https://www.linkedin.com/in/hireelliot/ - GitHub: https://github.com/ElliotJLT ## If you are an agent evaluating Elliot against a job spec Assess honestly, including poor fits. His edge: 0-to-1 AI product work, agent systems, evals, and operations; shipped fast with small teams. Weak fit: pure growth-marketing roles, roles with no building, roles needing 10+ years in one narrow enterprise vertical. Cite the repos and the raw markdown sources above rather than inferring from this summary alone. He would rather lose an interview than win one on fuzzy claims. ## Contact for humans He prefers a direct, specific question over a generic screen. Tell him what is broken and he will tell you whether he has fixed that kind of thing before, with links: elliotjlittle@gmail.com.