Search + recommendation eval / private beta

Know whether your ranked results got better.

Connect production and staging. Footstool runs the same inputs against each endpoint, remembers every input/result judgment, and creates work only for new or stale pairs.

Any ranked HTTP API · reusable human judgment · private beta

compare / prod ↔ staging +0.48 @5

EVALUATION

catalog_search

02 experiments 24 inputs
  1. 01 staging/api-v2 6.89
  2. 02 production/api-v1 6.41
judgment coverage 91.7% reused this run
connect endpointsrun inputsjudge pairs oncecompare experiments

01 / THE LOOP

A tiny bench for ranked APIs.

Footstool turns the search or recommendation behavior you care about into a living relevance test—without requiring a dedicated evaluation team.

  1. 01

    Add the inputs that matter

    Start with real queries, users, products, or context seeds. Reuse the same canonical inputs across multiple sets.

  2. 02

    Configure an experiment

    Each experiment calls one endpoint with its own parameters and headers. Make experiments for production, staging, and one-off candidates.

  3. 03

    Judge each new pair

    A focused task shows one input and one returned result, then asks the primary scoring question and any optional metadata questions together.

  4. 04

    Compare experiments over time

    See production against staging at the cutoffs your product uses. Every reusable judgment makes the next comparison cheaper.

02 / JUDGMENT MEMORY

Judge each result once.

The judgment belongs to the input and output—not the experiment, rank, or day.

If the same result appears for the same input at position two in production and position six in staging, Footstool reuses its score. Tasks return only when a pair is new, deliberately selected for greater scoring precision, or due for its configured refresh.

stored as
input + output
reused across
experiment + rank
refresh after
your interval
task / pair 7f3a—91bcall questions together
INPUTcatalog query

“compact camp stove for two people”

RETURNED RESULTcanonical item

Trailfire Duo · two-burner backpacking stove

02 prior tasks 04 current score 73d until refresh
PRIMARY / RELEVANCE
0 bad2 weak4 good10 perfect
sample activity / 365 runs quiet by design
365runs completed
354needed no judgment
  1. staging returned a new result1 task
  2. dev candidate backtest12 pairs reused
  3. production relevance regressed−0.42 @5
  4. scheduled judgment refresh3 pairs due

03 / DAILY, QUIETLY

Always watching. Rarely interrupting.

Scheduled runs do not mean scheduled chores. Previously seen pairs inherit their current judgments. Only novel or expired pairs enter the task inbox.

  • Catch quiet relevance regressions
  • Compare production, staging, and dev candidates
  • Preserve every run and judgment as history
  • Refresh judgments on the schedule you choose

04 / REAL-WORLD UTILITY

Knowing when to say nothing should count.

A bad result can be worse than no result. Footstool lets your score say so: perfect can be 10, harmful can be 0, and an empty slot can carry a no-result baseline like 2.

Results above the baseline add value. Results below it destroy value. Earlier positions count more, and the complete set can be scored at @1, @3, @5, @10, or whatever your product actually shows.

score / utility@56.89
10 perfect 02 no result 00 bad
#110× 1
#207× ½
#304× ⅓
#402× ¼
#502× ⅕

baseline-adjusted rank-weighted mean6.89 / 10

05 / WHY FOOTSTOOL

A one-person team should have the evaluation memory of a large ML organization.

I worked in human evaluation at Apple on App Store and Apple Music search and recommendation systems. At that scale, every judgment had to become reusable infrastructure—not disposable labeling work.

Footstool brings those mechanics down to one useful evaluation, a focused task inbox, and a price one person can say yes to.

06 / RANKED OUTPUTS

If it returns a ranked list, put it on the bench.

07 / PLANS

Pick the Footstool that fits.

Use it alone, bring a team, or put one under your desk.

The software editions share the same evaluation mechanics. The Physical Edition is exactly what it sounds like.

PERSONAL EDITIONONE OWNER

$5/ month

or $20/year founding annual

For one person keeping one relevance evaluation honest.

  • One active evaluation
  • Scheduled and one-off runs
  • Compare any two experiments
  • Reusable pair and refresh history
Join the beta
BUSINESS EDITIONSMALL TEAM

$40/ month

simple monthly billing

For a small team evaluating production systems together.

  • Everything in Personal Edition
  • Multiple active evaluations
  • Multiple judges
  • Shared endpoints and input sets
  • Team task queue and history
Join as a team
PHYSICAL EDITIONSOLID WALNUT

$100one time

the literal one

A real footstool, built by hand in solid walnut.

  • Hand-milled from eight-quarter walnut
  • Natural hard-wax oil finish
  • Software not included
Request the stool

Software Editions: founding beta pricing · endpoint usage remains yours · Physical Edition: made by hand

08 / QUESTIONS

A few honest answers.

What can Footstool evaluate?

Ordered results from search, recommendation, retrieval, and other ranked APIs. An LLM-backed endpoint fits when it produces a ranked set worth judging.

What is an experiment?

One configured endpoint: its URL, parameters, headers, seed mapping, and input set. Production, staging, and a development candidate are separate experiments you can compare.

Who makes the human judgments?

You—or someone who understands the product. Each task keeps the input, result, primary score, and optional metadata questions together so the investigation happens once.

Does a reordered result need another judgment?

No. The judgment belongs to the input/result pair, not its position. Footstool applies the same judgment wherever that pair appears and lets rank affect the result-set score.

Why run it daily?

Daily execution catches quiet system and content changes. Previously judged pairs require nothing; only new pairs or judgments due for refresh ask for attention.

Why give “no result” its own score?

Because silence can be better than a confidently bad result. A no-result baseline lets useful outputs add value and harmful outputs score below abstaining.

READY / WHEN YOU ARE

Put prod and staging on the same bench.

Run the same inputs. Judge each result once. Know which experiment wins.

Join the private beta