AI planner runs

48 runs over 48 cases · 277 turns · 8,299,837 tokens · ≈ 1.27 USD · 5 of 48 scored runs have a failing check

Scorecard

ClaimPassFail?
A pick that accepts the venue offer gets the search that turn 24 3 21
A capability is called by turn 2 35 1 12
A reply that asks puts the question on its final line 47 1
A search without an area is country-wide by design 48
An interpreted fact reaches the loop and not the account 5 43
A fact the couple declined is never asked again 10 38
Every figure carries a source it can support 10 38
A checklist built from a date is reported with its real task count 38 10
A year-only plan that talks timing says it is dated from June 6 7
Asking never outruns the two-in-a-row guard 43 5
A turn that writes but unlocks nothing uses the one-line form 43 5
A turn that changes the record shows a receipt 43 5
A terse message never gets an essay back 43 5
The focus chip never follows its own answer 43 5
A readiness offer is made at most twice per capability, then goes quiet 45 3
No three consecutive turns press for a fact the couple did not give 47 1
The interview detector sees every turn 47 1
No internal stage name is ever said to the couple 48
No word from inside the machine reaches the couple 48
A turn where the couple states nothing produces no receipt 48
The receipt reports a change, not a state 48
One turn builds the checklist at most once 48
The offer log keeps every real offer and at most the latest quiet turn 48
The opening rail holds at most four chips 47
Every chip speaks as the couple, never the singular 48
The card narrates the save; the prose never does 48
The run completed without a stream or write error 48
A year-only date gets no months-out figure 13
A year-only plan is not called late without saying why 13
"Next year" resolves to the calendar year after today 1

Every claim, counted across every scored run. ? means the run could not answer it, which is not a pass — a claim answered by nothing is untested, not met. Some passes are vacuous: a search that never ran cannot run without a search area.

Set against set

Claimv1
P / F / ?
v2
P / F / ?
v3
P / F / ?
No internal stage name is ever said to the couple 15 / 0 / 020 / 0 / 013 / 0 / 0
No word from inside the machine reaches the couple 15 / 0 / 020 / 0 / 013 / 0 / 0
No three consecutive turns press for a fact the couple did not give 14 / 0 / 120 / 0 / 013 / 0 / 0
Asking never outruns the two-in-a-row guard 11 / 0 / 419 / 0 / 113 / 0 / 0
The interview detector sees every turn 14 / 0 / 120 / 0 / 013 / 0 / 0
A turn where the couple states nothing produces no receipt 15 / 0 / 020 / 0 / 013 / 0 / 0
A turn that writes but unlocks nothing uses the one-line form 12 / 0 / 320 / 0 / 011 / 0 / 2
The receipt reports a change, not a state 15 / 0 / 020 / 0 / 013 / 0 / 0
A checklist built from a date is reported with its real task count 9 / 0 / 620 / 0 / 09 / 0 / 4
One turn builds the checklist at most once 15 / 0 / 020 / 0 / 013 / 0 / 0
A fact the couple declined is never asked again 6 / 0 / 92 / 0 / 182 / 0 / 11
A readiness offer is made at most twice per capability, then goes quiet 14 / 0 / 118 / 0 / 213 / 0 / 0
The offer log keeps every real offer and at most the latest quiet turn 15 / 0 / 020 / 0 / 013 / 0 / 0
A search without an area is country-wide by design 0 / 0 / 150 / 0 / 200 / 0 / 13
The opening rail holds at most four chips 15 / 0 / 019 / 0 / 013 / 0 / 0
Every chip speaks as the couple, never the singular 15 / 0 / 020 / 0 / 013 / 0 / 0
Every figure carries a source it can support 2 / 0 / 135 / 0 / 153 / 0 / 10
An interpreted fact reaches the loop and not the account 1 / 0 / 143 / 0 / 171 / 0 / 12
The card narrates the save; the prose never does 15 / 0 / 020 / 0 / 013 / 0 / 0
A turn that changes the record shows a receipt 12 / 0 / 320 / 0 / 011 / 0 / 2
A terse message never gets an essay back 13 / 0 / 219 / 0 / 111 / 0 / 2
The focus chip never follows its own answer 12 / 0 / 320 / 0 / 011 / 0 / 2
The run completed without a stream or write error 15 / 0 / 020 / 0 / 013 / 0 / 0
A year-only date gets no months-out figure 5 / 0 / 07 / 0 / 01 / 0 / 0
A year-only plan that talks timing says it is dated from June 2 / 0 / 34 / 0 / 30 / 0 / 1
A year-only plan is not called late without saying why 5 / 0 / 07 / 0 / 01 / 0 / 0
"Next year" resolves to the calendar year after today 1 / 0 / 0
A pick that accepts the venue offer gets the search that turn 5 / 0 / 1011 / 2 / 78 / 1 / 4
A capability is called by turn 2 10 / 0 / 516 / 1 / 39 / 0 / 4
A reply that asks puts the question on its final line 14 / 1 / 020 / 0 / 013 / 0 / 0

The same claims, split by which generation of cases produced them: v1 15 runs, 1 with a failing check · v2 20 runs, 3 with a failing check · v3 13 runs, 1 with a failing check. Ordered worst gap first: a claim at the top holds in the older sets and breaks in the newest one, which is how a rule that fits the cases looks from the outside. Sets are not comparable case for case — a claim unanswered in one set and answered in another means the cases differ, not that the coach changed.

Runs, newest first

focus-echo v3 4 turns ≈ 0.015 USD all checks pass 3 unanswered

What to focus on, asked twice in a row after full facts. Scores whether the focus suggestion recurs after being answered, and whether the two answers differ.

2026-08-20 02:17 · commit bad81caeb34 · 124,763 tokens · 1 receipt · tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice
full-journey v1 15 turns ≈ 0.102 USD 1 fail 4 unanswered

Complete onboarding then venue work, photographer costs, a reload and a month plan. Scores cadence and state over a 15-step conversation.

2026-08-20 02:14 · commit bad81caeb34 · 687,899 tokens · 4 receipts · tools: executeClientActions, getBudgetBreakdown, searchWeddingSuppliers, getCoachReference, getUpcomingTaskAdvice, planningSnapshot
1 failing check
  • A reply that asks puts the question on its final line
    turns 7 buried the question
changing-minds v1 14 turns ≈ 0.095 USD all checks pass 3 unanswered

Full facts at once, then a guest cut, a date move, a city switch and re-searches. Scores writes, rebuilds and re-ranking over 14 steps.

2026-08-20 02:11 · commit bad81caeb34 · 666,551 tokens · 4 receipts · tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown, getUpcomingTaskAdvice, planningSnapshot
returning-couple v1 5 turns ≈ 0.011 USD all checks pass 6 unanswered

A first meeting, then a return on another day with the transcript gone. Scores whether the second opener knows them, and whether it re-asks what it was told.

2026-08-20 02:10 · commit bad81caeb34 · 64,610 tokens · 1 receipt · tools: executeClientActions
year-then-month v3 6 turns ≈ 0.027 USD all checks pass 4 unanswered

A bare 2028, a timeline question, then September 2028. Scores the June anchor, the rebuild, and whether any internal vocabulary reaches the re-dating reply.

2026-08-20 02:04 · commit bad81caeb34 · 191,965 tokens · 3 receipts · tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice
locked-door-return v3 6 turns ≈ 0.021 USD all checks pass 6 unanswered

A date declined twice, a quiet stretch, then "what would our timeline look like?". Scores whether a couple asking for a locked capability is told its price on a quiet turn.

2026-08-20 02:02 · commit bad81caeb34 · 117,483 tokens · 1 receipt · tools: getBudgetBreakdown, getCoachReference, executeClientActions
show-your-workings v3 5 turns ≈ 0.014 USD all checks pass 8 unanswered

A cost question, then two challenges to the figure. Scores whether source detail escalates under challenge instead of appearing everywhere or nowhere.

2026-08-20 02:00 · commit bad81caeb34 · 146,331 tokens · 0 receipts · tools: getBudgetBreakdown, getCoachReference, readCoachArticle
just-engaged v3 4 turns ≈ 0.0091 USD all checks pass 10 unanswered

Engaged last night, knows nothing, wants nothing yet. Scores whether the coach meets the excitement before reaching for planning.

2026-08-20 01:59 · commit bad81caeb34 · 54,849 tokens · 0 receipts · tools: none called
man-of-few-words v3 5 turns ≈ 0.020 USD all checks pass 4 unanswered

A couple who answers in fragments and never elaborates. Scores whether reply length and question pressure come down to match.

2026-08-20 01:57 · commit bad81caeb34 · 126,583 tokens · 1 receipt · tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice, planningSnapshot
booked-or-browsing v3 5 turns ≈ 0.024 USD all checks pass 4 unanswered

A place and a date given with no venue intent, and the venue turns out to be booked. Scores whether the coach checks before assuming a venue search.

2026-08-20 01:56 · commit bad81caeb34 · 145,223 tokens · 2 receipts · tools: executeClientActions, searchWeddingSuppliers
broad-then-pick v3 6 turns ≈ 0.028 USD all checks pass 5 unanswered

A venue search on location and guests only, a request to pick, then budget and a priority, then the same request. Scores whether recommendation strength follows what is known.

2026-08-20 01:55 · commit bad81caeb34 · 160,492 tokens · 2 receipts · tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown
style-not-place v3 6 turns ≈ 0.035 USD 1 fail 4 unanswered

A coastal-wedding wish given before any geography. Scores what a setting phrase writes, and whether the coach still asks plainly where.

2026-08-20 01:54 · commit bad81caeb34 · 238,665 tokens · 2 receipts · tools: searchWeddingKnowledge, executeClientActions, searchWeddingSuppliers
1 failing check
  • A pick that accepts the venue offer gets the search that turn
    turns 5 only confirmed the save
deadline-stonewall v3 6 turns ≈ 0.026 USD all checks pass 4 unanswered

Ten weeks out, and the couple refuses questions. Scores whether the ask guard holds when urgency and a stonewall land together.

2026-08-20 01:52 · commit bad81caeb34 · 142,101 tokens · 2 receipts · tools: getCoachReference, executeClientActions, searchWeddingSuppliers
card-left-open v3 5 turns ≈ 0.019 USD all checks pass 5 unanswered

A county named, the location card left open, and the conversation carries on. Scores what the coach says about an unpicked place, and whether the later pick lands.

2026-08-20 01:51 · commit bad81caeb34 · 109,917 tokens · 2 receipts · tools: executeClientActions, searchWeddingSuppliers
all-in-one-line v3 4 turns ≈ 0.015 USD all checks pass 4 unanswered

Date, guests, budget and place in a single message, then "did you get all that?". Scores the receipt contract on a multi-write turn and the reply that follows it.

2026-08-20 01:50 · commit bad81caeb34 · 107,262 tokens · 1 receipt · tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice
note-taker v3 5 turns ≈ 0.018 USD all checks pass 6 unanswered

Facts given conversationally, then the couple asks if any of it was written down. Scores save narration in prose against receipt cards that must exist.

2026-08-20 01:49 · commit bad81caeb34 · 95,620 tokens · 2 receipts · tools: executeClientActions
two-brides v2 6 turns ≈ 0.028 USD all checks pass 3 unanswered

Two brides who give both names. Scores how the couple is addressed, and whether a same-sex request is answered with suppliers rather than reassurance.

2026-08-20 01:47 · commit bad81caeb34 · 192,689 tokens · 3 receipts · tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown, getUpcomingTaskAdvice
budget-first v2 6 turns ≈ 0.026 USD all checks pass 5 unanswered

The budget given first and the date last. Scores what is offered when the money is known and nothing else, and whether the ranking follows the facts.

2026-08-20 01:46 · commit bad81caeb34 · 123,323 tokens · 3 receipts · tools: executeClientActions
rail-first v2 6 turns ≈ 0.029 USD all checks pass 5 unanswered

The conversation opened from a rail chip rather than the opening button, then driven by chips. Scores the chip entry point and what a tapped turn writes.

2026-08-20 01:45 · commit bad81caeb34 · 216,447 tokens · 2 receipts · tools: executeClientActions, getUpcomingTaskAdvice, searchWeddingKnowledge, searchWeddingSuppliers, compareSuppliers
pick-then-reload v2 6 turns ≈ 0.024 USD all checks pass 4 unanswered

A location picked, then an immediate reload. Scores whether the pick survived the tab, and whether the search that was owed still arrives.

2026-08-20 01:43 · commit bad81caeb34 · 149,404 tokens · 2 receipts · tools: executeClientActions, searchWeddingSuppliers
budget-reality v2 6 turns ≈ 0.028 USD 1 fail 5 unanswered

£8,000 for 120 guests in London. Scores whether the gap is named plainly with grounded figures, and what the search returns against it.

2026-08-20 01:42 · commit bad81caeb34 · 160,251 tokens · 3 receipts · tools: executeClientActions, getBudgetBreakdown, getCoachReference, searchWeddingSuppliers
1 failing check
  • A pick that accepts the venue offer gets the search that turn
    turns 3 only confirmed the save
partners-disagree v2 7 turns ≈ 0.036 USD all checks pass 6 unanswered

A barn against a city hotel, a band against a DJ, and a request to settle it. Scores whether an opinion is given, and what a stated preference writes.

2026-08-20 01:40 · commit bad81caeb34 · 226,783 tokens · 1 receipt · tools: getCoachReference, executeClientActions, getBudgetBreakdown
month-plan v2 6 turns ≈ 0.028 USD all checks pass 4 unanswered

A date, then three questions about the checklist it built. Scores task advice against a real list, and whether nothing is invented for a month that has no tasks.

2026-08-20 01:39 · commit bad81caeb34 · 201,421 tokens · 2 receipts · tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice
shortlist-compare v2 5 turns ≈ 0.019 USD 1 fail 5 unanswered

A venue search, then the compare chip it offers, then a recommendation. Scores the compare capability and whether the choice is argued from the couple’s own facts.

2026-08-20 01:38 · commit bad81caeb34 · 124,915 tokens · 1 receipt · tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown, compareSuppliers
1 failing check
  • A pick that accepts the venue offer gets the search that turn
    turns 2 only confirmed the save
named-venue v2 5 turns ≈ 0.026 USD all checks pass 4 unanswered

A specific venue named by a friend, then a comparison. Scores the by-name lookup and whether an unknown name is admitted rather than described.

2026-08-20 01:37 · commit bad81caeb34 · 179,670 tokens · 1 receipt · tools: executeClientActions, searchWeddingSuppliers, searchSupplierByName, getCoachReference
off-peak-thursday v2 6 turns ≈ 0.033 USD all checks pass 4 unanswered

A Thursday in February 2028, picked to save money. Scores the date write, the saving claim, and whether the season is answered from a source.

2026-08-20 01:35 · commit bad81caeb34 · 210,620 tokens · 3 receipts · tools: executeClientActions, getCoachReference, getBudgetBreakdown, searchWeddingSuppliers
money-no-object v2 5 turns ≈ 0.020 USD 1 fail 3 unanswered

A budget described as unlimited and never numbered. Scores whether a figure is written for a couple who gave none.

2026-08-20 01:34 · commit bad81caeb34 · 123,690 tokens · 1 receipt · tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown
1 failing check
  • A capability is called by turn 2
    first at turn 3 (executeClientActions, searchWeddingSuppliers)
big-wedding v2 6 turns ≈ 0.032 USD all checks pass 3 unanswered

Three hundred guests and £120,000. Scores the figures and the search at a size most weddings never reach.

2026-08-20 01:33 · commit bad81caeb34 · 218,672 tokens · 1 receipt · tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown, getCoachReference, getUpcomingTaskAdvice
micro-wedding v2 6 turns ≈ 0.027 USD all checks pass 5 unanswered

Fourteen guests and £6,000. Scores whether the figures and the plan scale down, or whether a small wedding is answered with big-wedding numbers.

2026-08-20 01:31 · commit bad81caeb34 · 139,310 tokens · 3 receipts · tools: executeClientActions, getBudgetBreakdown
guest-list-politics v2 6 turns ≈ 0.028 USD all checks pass 5 unanswered

Forty extra guests demanded against a fixed budget. Scores whether the facts inside a complaint get read, and whether the answer is advice or another question.

2026-08-20 01:30 · commit bad81caeb34 · 170,167 tokens · 2 receipts · tools: getCoachReference, executeClientActions, getBudgetBreakdown, getUpcomingTaskAdvice
sceptical-couple v2 6 turns ≈ 0.017 USD all checks pass 6 unanswered

Three challenges to the coach itself before any planning. Scores honesty about what it is, and whether pressure makes it leak its own vocabulary.

2026-08-20 01:29 · commit bad81caeb34 · 133,304 tokens · 1 receipt · tools: getCoachReference, executeClientActions, searchWeddingSuppliers
messy-typing v2 6 turns ≈ 0.029 USD all checks pass 4 unanswered

Every fact given in typos and abbreviations. Scores what extraction writes from unpunctuated input, and whether it invents what it could not read.

2026-08-20 01:27 · commit bad81caeb34 · 182,436 tokens · 3 receipts · tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice, getCoachReference
self-contradiction v2 7 turns ≈ 0.036 USD all checks pass 5 unanswered

A date, a guest count and a budget, each overwritten a turn later. Scores which value survives, on the receipt and in the profile.

2026-08-20 01:25 · commit bad81caeb34 · 188,474 tokens · 4 receipts · tools: executeClientActions, planningSnapshot
already-booked v2 6 turns ≈ 0.044 USD all checks pass 3 unanswered

Venue and photographer already booked. Scores whether the coach re-offers what is done, and what it offers instead.

2026-08-20 01:24 · commit bad81caeb34 · 351,282 tokens · 1 receipt · tools: executeClientActions, searchWeddingSuppliers, planningSnapshot, getUpcomingTaskAdvice, getBudgetBreakdown, getCoachReference
imminent-wedding v2 5 turns ≈ 0.021 USD all checks pass 4 unanswered

A wedding six weeks away. Scores the plan for a date with no runway, and whether the coach onboards a couple who is out of time.

2026-08-20 01:23 · commit bad81caeb34 · 140,145 tokens · 2 receipts · tools: executeClientActions, getBudgetBreakdown, getUpcomingTaskAdvice, getCoachReference
abroad-wedding v2 6 turns ≈ 0.029 USD all checks pass 4 unanswered

A destination wedding in Puglia, then a request to search there. Scores what a non-UK place writes, and whether the coach searches anyway or says what it cannot do.

2026-08-20 01:21 · commit bad81caeb34 · 176,844 tokens · 2 receipts · tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown, getCoachReference, getUpcomingTaskAdvice
budget-journey v1 12 turns ≈ 0.077 USD all checks pass 5 unanswered

A budget refusal that later becomes £18,000, then splits, benchmarks and savings questions at their own numbers. 12 steps.

2026-08-20 01:15 · commit bad81caeb34 · 510,455 tokens · 4 receipts · tools: executeClientActions, getCoachReference, searchWeddingSuppliers, getBudgetBreakdown, getUpcomingTaskAdvice
relative-date v1 3 turns ≈ 0.0087 USD all checks pass 7 unanswered

The couple says "next year" and never names it. Scores whether the written year is the calendar year after today, on the receipt and in the profile.

2026-08-20 01:11 · commit bad81caeb34 · 65,567 tokens · 1 receipt · tools: executeClientActions, getUpcomingTaskAdvice
cost-questions v1 4 turns ≈ 0.012 USD all checks pass 7 unanswered

Three direct cost questions. Scores whether every figure is grounded and whether any is attributed to a source that does not hold it.

2026-08-20 01:10 · commit bad81caeb34 · 87,715 tokens · 1 receipt · tools: getBudgetBreakdown, getCoachReference
supplier-search v1 4 turns ≈ 0.016 USD all checks pass 5 unanswered

A real place picked from the location field, then a venue request. Scores whether the pick set a search area and unlocked supplier search, and what the search returned.

2026-08-20 01:10 · commit bad81caeb34 · 132,374 tokens · 1 receipt · tools: executeClientActions, searchWeddingSuppliers, compareSuppliers
stonewaller v1 6 turns ≈ 0.018 USD all checks pass 9 unanswered

Five turns that state nothing at all. Scores the ask density over a whole conversation, and whether any turn without a fact still delivered value.

2026-08-20 01:08 · commit bad81caeb34 · 100,169 tokens · 0 receipts · tools: getCoachReference
just-show-me v1 4 turns ≈ 0.016 USD all checks pass 5 unanswered

A demand for suppliers before anything is known, repeated. Scores whether an explicit request is ever refused outright, and what is delivered instead.

2026-08-20 01:07 · commit bad81caeb34 · 112,382 tokens · 1 receipt · tools: searchWeddingSuppliers, executeClientActions
refusal-reload v1 5 turns ≈ 0.0097 USD all checks pass 8 unanswered

A budget refusal, then the tab closes and reopens. Scores whether the deferral survived the reload and whether the coach asks again after it.

2026-08-20 01:06 · commit bad81caeb34 · 55,546 tokens · 0 receipts · tools: none called
generic-location v1 4 turns ≈ 0.016 USD all checks pass 6 unanswered

A location too generic to search on. Scores whether supplier search stays locked, and whether the coach names the missing key instead of searching anyway.

2026-08-20 01:05 · commit bad81caeb34 · 96,596 tokens · 1 receipt · tools: executeClientActions, searchWeddingSuppliers
all-at-once v1 4 turns ≈ 0.017 USD all checks pass 4 unanswered

Every core fact in two messages. Scores the full receipt card, the unlocks it names, and whether the next turn re-reports facts it did not change.

2026-08-20 01:04 · commit bad81caeb34 · 115,414 tokens · 2 receipts · tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice
vision-first v1 6 turns ≈ 0.023 USD all checks pass 4 unanswered

Vision, then a guest count, then two budget refusals, then a year. Scores what a rapport fact writes, and whether a refusal is honoured.

2026-08-20 01:03 · commit bad81caeb34 · 125,735 tokens · 2 receipts · tools: executeClientActions, getUpcomingTaskAdvice
year-only v1 5 turns ≈ 0.018 USD all checks pass 5 unanswered

Two budget refusals, then the bare year 2027. Scores the year-only plan, the receipt for a checklist build, and whether a declined fact is asked again.

2026-08-20 01:01 · commit bad81caeb34 · 96,574 tokens · 1 receipt · tools: executeClientActions, getUpcomingTaskAdvice
opener v1 1 turn ≈ 0.0023 USD all checks pass 15 unanswered

The opening move alone. Scores the tier-0 insight, the provenance clause, and whether the single question is an offer’s price rather than an interview opener.

2026-08-20 01:01 · commit bad81caeb34 · 11,149 tokens · 0 receipts · tools: none called