“compact camp stove for two people”
Search + recommendation eval / private beta
Know whether your ranked results got better.
Connect production and staging. Footstool runs the same inputs against each endpoint, remembers every input/result judgment, and creates work only for new or stale pairs.
Any ranked HTTP API · reusable human judgment · private beta
EVALUATION
catalog_search
- 01 staging/api-v2 6.89
- 02 production/api-v1 6.41
01 / THE LOOP
A tiny bench for ranked APIs.
Footstool turns the search or recommendation behavior you care about into a living relevance test—without requiring a dedicated evaluation team.
-
01
Add the inputs that matter
Start with real queries, users, products, or context seeds. Reuse the same canonical inputs across multiple sets.
-
02
Configure an experiment
Each experiment calls one endpoint with its own parameters and headers. Make experiments for production, staging, and one-off candidates.
-
03
Judge each new pair
A focused task shows one input and one returned result, then asks the primary scoring question and any optional metadata questions together.
-
04
Compare experiments over time
See production against staging at the cutoffs your product uses. Every reusable judgment makes the next comparison cheaper.
02 / JUDGMENT MEMORY
Judge each result once.
The judgment belongs to the input and output—not the experiment, rank, or day.
If the same result appears for the same input at position two in production and position six in staging, Footstool reuses its score. Tasks return only when a pair is new, deliberately selected for greater scoring precision, or due for its configured refresh.
- stored as
- input + output
- reused across
- experiment + rank
- refresh after
- your interval
Trailfire Duo · two-burner backpacking stove
- staging returned a new result1 task
- dev candidate backtest12 pairs reused
- production relevance regressed−0.42 @5
- scheduled judgment refresh3 pairs due
03 / DAILY, QUIETLY
Always watching. Rarely interrupting.
Scheduled runs do not mean scheduled chores. Previously seen pairs inherit their current judgments. Only novel or expired pairs enter the task inbox.
- Catch quiet relevance regressions
- Compare production, staging, and dev candidates
- Preserve every run and judgment as history
- Refresh judgments on the schedule you choose
04 / REAL-WORLD UTILITY
Knowing when to say nothing should count.
A bad result can be worse than no result. Footstool lets your score say so: perfect can be 10, harmful can be 0, and an empty slot can carry a no-result baseline like 2.
Results above the baseline add value. Results below it destroy value. Earlier positions count more, and the complete set can be scored at @1, @3, @5, @10, or whatever your product actually shows.
baseline-adjusted rank-weighted mean6.89 / 10
05 / WHY FOOTSTOOL
A one-person team should have the evaluation memory of a large ML organization.
I worked in human evaluation at Apple on App Store and Apple Music search and recommendation systems. At that scale, every judgment had to become reusable infrastructure—not disposable labeling work.
Footstool brings those mechanics down to one useful evaluation, a focused task inbox, and a price one person can say yes to.
06 / RANKED OUTPUTS
If it returns a ranked list, put it on the bench.
- 01
Did staging improve search for the long-tail queries we care about?
- 02
Did the new recommender put the useful candidates first?
- 03
Does retrieval surface good evidence before noisy context?
- 04
Would returning nothing be better than the result we show today?
07 / PLANS
Pick the Footstool that fits.
Use it alone, bring a team, or put one under your desk.
The software editions share the same evaluation mechanics. The Physical Edition is exactly what it sounds like.
$5/ month
or $20/year founding annualFor one person keeping one relevance evaluation honest.
- One active evaluation
- Scheduled and one-off runs
- Compare any two experiments
- Reusable pair and refresh history
$40/ month
simple monthly billingFor a small team evaluating production systems together.
- Everything in Personal Edition
- Multiple active evaluations
- Multiple judges
- Shared endpoints and input sets
- Team task queue and history
$100one time
the literal oneA real footstool, built by hand in solid walnut.
- Hand-milled from eight-quarter walnut
- Natural hard-wax oil finish
- Software not included
Software Editions: founding beta pricing · endpoint usage remains yours · Physical Edition: made by hand
08 / QUESTIONS
A few honest answers.
What can Footstool evaluate?
Ordered results from search, recommendation, retrieval, and other ranked APIs. An LLM-backed endpoint fits when it produces a ranked set worth judging.
What is an experiment?
One configured endpoint: its URL, parameters, headers, seed mapping, and input set. Production, staging, and a development candidate are separate experiments you can compare.
Who makes the human judgments?
You—or someone who understands the product. Each task keeps the input, result, primary score, and optional metadata questions together so the investigation happens once.
Does a reordered result need another judgment?
No. The judgment belongs to the input/result pair, not its position. Footstool applies the same judgment wherever that pair appears and lets rank affect the result-set score.
Why run it daily?
Daily execution catches quiet system and content changes. Previously judged pairs require nothing; only new pairs or judgments due for refresh ask for attention.
Why give “no result” its own score?
Because silence can be better than a confidently bad result. A no-result baseline lets useful outputs add value and harmful outputs score below abstaining.
READY / WHEN YOU ARE
Put prod and staging on the same bench.
Run the same inputs. Judge each result once. Know which experiment wins.
Join the private beta