Built
I build the part that decides whether output is good enough to ship. Usually somewhere a wrong answer costs a student a grade, or misses what a child was trying to tell someone.
the through-line
Producing plausible output is no longer the hard part. Knowing whether to ship it still is, and that gap is where most of my work has gone. The same question keeps surfacing at different points in a system: what got checked, what got rejected, and who decided. Before a student sees a mark. Before a child's disclosure gets missed. Before an application goes out. Before an agent changes this page.
It applies to the model and to the person equally. You read what the loop made, and you can defend what carries your name. crux is the instrument I built to find out whether that actually holds.
- Zero Gravity AI STEM tutor
An always-on evaluator grades every coaching session against the Socratic spec, and marking is tested against official mark schemes before a student sees a grade. That eval infrastructure took marking accuracy from a ~67% bare-model baseline to over 99%. Live across four STEM subjects on every major UK exam board, first commit to App Store in 45 days, and it was selected as one of eight companies nationally for the DfE and DSIT AI Tutoring Tools Pioneers Programme, which requires meeting the government's Generative AI Product Safety Standards.
- ward
Decides which messages from a child are genuine safeguarding disclosures and routes those to a named human on a clock, grounded in KCSIE rather than keyword matching. Built around precision, because a Designated Safeguarding Lead who gets paged on every false alarm learns to ignore the alerts, which is worse than having none. On its published synthetic sets the Claude judge reaches 90% recall at 100% precision, against 50/83 for a keyword baseline.
- boulot
Three agents with opposing briefs, a hiring manager, a reviewer and a strategist, argue over a CV before it is allowed out. I ran my own search through it, then open-sourced it. My partner and my sister use it too.
- crux
The judgement no commit log records: what a person rejected, redirected, or killed while the model did the typing. Ongoing research, with the method, the results run on myself, the honest objections and the limitations all published.
- this site
An agent proposes one change, a rubric scores it against explicit criteria, and a human merges it or does not. Every cadence, gate and stopping rule is published on /loops, and the spend is metered.
track record
Roles and the outcome that mattered in each. Full history, titles, and references on LinkedIn.
Built and deployed a production multi-agent AI STEM tutor. First commit to App Store in 45 days; marking accuracy raised from a ~67% baseline to over 99% against official mark schemes; selected as one of eight companies nationally for the DfE and DSIT AI Tutoring Tools Pioneers Programme, safe AI tutoring for disadvantaged pupils.
- Farewill
Product work in an SRA/FCA-regulated environment.
- MealsForTheNHS
Co-founded a pandemic-response effort that raised £1.8m and delivered 303,000 meals to NHS staff.
- Flash Pack
Founding-team product work; helped take the company from pre-seed to Series A.
reference
If I were starting another business tomorrow, Elliot would be one of the first people I'd try to hire.
He doesn't just understand what AI makes possible; he builds with it, experiments rapidly and ships. He now contributes as many PRs as some of our senior engineers, while bringing the customer insight and commercial judgement you would expect from a strong product leader.
Elliot drove the development of Zero Gravity Tutor, taking it from an idea to a product that has unlocked an exciting new opportunity for Zero Gravity.
in production
Zero Gravity AI STEM tutor
A private tutor at the shoulder of students whose families could not pay for one: live across four STEM subjects on every major UK exam board, first commit to the App Store in 45 days. Socratic by design, it coaches a student to the answer and will not hand it over, however creatively they ask, and an always-on evaluator grades every session against that spec. Coaching, practice, marking and assignments each run as their own agent with their own pedagogy and guardrails, and marking is tested against real past papers and official mark schemes: a ~67% bare-model baseline to over 99%. Selected as one of eight companies nationally for the DfE and DSIT AI Tutoring Tools Pioneers Programme, built to the government's Generative AI Product Safety Standards. I built and deployed it at Zero Gravity.
research
crux
You are shipping faster than ever. Are you getting sharper, or just getting carried? Nothing currently measures that. I noticed it in myself at Zero Gravity: shipping faster than I ever had, and slower to say what I would have done differently. Output has never been higher and no instrument tells you whether the person behind it is improving, plateauing or quietly atrophying. Fluency frameworks answer the baseline and everyone will pass them; the layer above is where the difference sits, in trust calibration, resistance to output that looks polished, and knowing what to kill.
crux measures it. A Claude Code hook reads each session and extracts what the human actually decided: what got rejected, redirected or killed while the model did the typing. Ongoing research rather than a product, published with the method, the results run on myself, the honest objections, the limitations, and a memo to the platform layer about the half that nothing measures.
agent tools
boulot
Open-source career-ops system that runs on your own laptop through Claude Code. Tailors your CV per role, then three adversarial agents (hiring manager, reviewer, strategist) fight over the draft. Built for my own search in a brutal market; it worked, so I open-sourced it.
claude-skill-potions
Curated Claude Code skills for ops and product workflows. The skills directory is 40k+ deep; these are the ones that actually work.
vox
Voice of Customer research agent. Eight days of PM research in eight minutes: JTBD, personas, opportunity mapping from Gong, Granola and Jiminy data.
dabble
Visual editor for server-rendered (Hotwire) apps. Edit the running app in place, write real ERB. Kills the design-to-code handoff for the stacks React-first tools ignore.
ward
A safeguarding layer for LLM apps that serve under-18s. It screens each message for a genuine disclosure, separates that from ordinary bad conduct, and routes the real ones to a named human on a clock. Grounded in KCSIE rather than generic content moderation. Built around the precision problem: page a Designated Safeguarding Lead on every false alarm and they stop trusting the alerts, which is worse than having none. On its published synthetic eval sets, the Claude judge reaches 90% recall at 100% precision and a 0% false-positive rate, against a keyword baseline at 50/83/8.6.
homebuyer-mcp
UK home-buying MCP server. Conveyancers and mortgage brokers from live SRA, FCA and Companies House registers, plus stamp duty, lease checks, survey explainers and title register analysis. Eleven tools.
hooksmith
Browse and install pre-built Claude Code hooks with one command. Twelve hooks, zero config. The missing package manager for hooks.
everything else
- elliot-osMy site, run like a product: live projects, a public roadmap, a changelog, and pages maintained by agents.
- lamplightOne day project. See which streets are lit before you run in the dark.
- spawn-cafe ★1Share your link, pick a slot, find a cafe halfway. Agentic coffee meetup scheduler.
- mo-hanzi墨 mò — menubar SRS for learning Chinese characters (HSK 1-3)
- claudemd-lintLint your CLAUDE.md files. Catch vague rules, bloated configs, and instructions that should be hooks.
- zestemacOS menubar companion for Claude Code skills. Discover, install, and manage skills in seconds.
- spotifyunwrappedYour Spotify data, visualised properly. Privacy-first listening analytics that run entirely in your browser.
ideas or feedback?
Two systems here are real but private: argus, an agent fleet that reads the AI news every morning and writes me a brief, and LifeOS, the front door that routes my whole setup. Ask me about either: elliotjlittle@gmail.com.