Scorecard
| Claim | Pass | Fail | ? | |
|---|---|---|---|---|
| A pick that accepts the venue offer gets the search that turn | 24 | 3 | 21 | |
| A capability is called by turn 2 | 35 | 1 | 12 | |
| A reply that asks puts the question on its final line | 47 | 1 | ||
| A search without an area is country-wide by design | 48 | |||
| An interpreted fact reaches the loop and not the account | 5 | 43 | ||
| A fact the couple declined is never asked again | 10 | 38 | ||
| Every figure carries a source it can support | 10 | 38 | ||
| A checklist built from a date is reported with its real task count | 38 | 10 | ||
| A year-only plan that talks timing says it is dated from June | 6 | 7 | ||
| Asking never outruns the two-in-a-row guard | 43 | 5 | ||
| A turn that writes but unlocks nothing uses the one-line form | 43 | 5 | ||
| A turn that changes the record shows a receipt | 43 | 5 | ||
| A terse message never gets an essay back | 43 | 5 | ||
| The focus chip never follows its own answer | 43 | 5 | ||
| A readiness offer is made at most twice per capability, then goes quiet | 45 | 3 | ||
| No three consecutive turns press for a fact the couple did not give | 47 | 1 | ||
| The interview detector sees every turn | 47 | 1 | ||
| No internal stage name is ever said to the couple | 48 | |||
| No word from inside the machine reaches the couple | 48 | |||
| A turn where the couple states nothing produces no receipt | 48 | |||
| The receipt reports a change, not a state | 48 | |||
| One turn builds the checklist at most once | 48 | |||
| The offer log keeps every real offer and at most the latest quiet turn | 48 | |||
| The opening rail holds at most four chips | 47 | |||
| Every chip speaks as the couple, never the singular | 48 | |||
| The card narrates the save; the prose never does | 48 | |||
| The run completed without a stream or write error | 48 | |||
| A year-only date gets no months-out figure | 13 | |||
| A year-only plan is not called late without saying why | 13 | |||
| "Next year" resolves to the calendar year after today | 1 |
Every claim, counted across every scored run. ? means the run could not answer it, which is not a pass — a claim answered by nothing is untested, not met. Some passes are vacuous: a search that never ran cannot run without a search area.
Set against set
| Claim | v1 P / F / ? | v2 P / F / ? | v3 P / F / ? |
|---|---|---|---|
| No internal stage name is ever said to the couple | 15 / 0 / 0 | 20 / 0 / 0 | 13 / 0 / 0 |
| No word from inside the machine reaches the couple | 15 / 0 / 0 | 20 / 0 / 0 | 13 / 0 / 0 |
| No three consecutive turns press for a fact the couple did not give | 14 / 0 / 1 | 20 / 0 / 0 | 13 / 0 / 0 |
| Asking never outruns the two-in-a-row guard | 11 / 0 / 4 | 19 / 0 / 1 | 13 / 0 / 0 |
| The interview detector sees every turn | 14 / 0 / 1 | 20 / 0 / 0 | 13 / 0 / 0 |
| A turn where the couple states nothing produces no receipt | 15 / 0 / 0 | 20 / 0 / 0 | 13 / 0 / 0 |
| A turn that writes but unlocks nothing uses the one-line form | 12 / 0 / 3 | 20 / 0 / 0 | 11 / 0 / 2 |
| The receipt reports a change, not a state | 15 / 0 / 0 | 20 / 0 / 0 | 13 / 0 / 0 |
| A checklist built from a date is reported with its real task count | 9 / 0 / 6 | 20 / 0 / 0 | 9 / 0 / 4 |
| One turn builds the checklist at most once | 15 / 0 / 0 | 20 / 0 / 0 | 13 / 0 / 0 |
| A fact the couple declined is never asked again | 6 / 0 / 9 | 2 / 0 / 18 | 2 / 0 / 11 |
| A readiness offer is made at most twice per capability, then goes quiet | 14 / 0 / 1 | 18 / 0 / 2 | 13 / 0 / 0 |
| The offer log keeps every real offer and at most the latest quiet turn | 15 / 0 / 0 | 20 / 0 / 0 | 13 / 0 / 0 |
| A search without an area is country-wide by design | 0 / 0 / 15 | 0 / 0 / 20 | 0 / 0 / 13 |
| The opening rail holds at most four chips | 15 / 0 / 0 | 19 / 0 / 0 | 13 / 0 / 0 |
| Every chip speaks as the couple, never the singular | 15 / 0 / 0 | 20 / 0 / 0 | 13 / 0 / 0 |
| Every figure carries a source it can support | 2 / 0 / 13 | 5 / 0 / 15 | 3 / 0 / 10 |
| An interpreted fact reaches the loop and not the account | 1 / 0 / 14 | 3 / 0 / 17 | 1 / 0 / 12 |
| The card narrates the save; the prose never does | 15 / 0 / 0 | 20 / 0 / 0 | 13 / 0 / 0 |
| A turn that changes the record shows a receipt | 12 / 0 / 3 | 20 / 0 / 0 | 11 / 0 / 2 |
| A terse message never gets an essay back | 13 / 0 / 2 | 19 / 0 / 1 | 11 / 0 / 2 |
| The focus chip never follows its own answer | 12 / 0 / 3 | 20 / 0 / 0 | 11 / 0 / 2 |
| The run completed without a stream or write error | 15 / 0 / 0 | 20 / 0 / 0 | 13 / 0 / 0 |
| A year-only date gets no months-out figure | 5 / 0 / 0 | 7 / 0 / 0 | 1 / 0 / 0 |
| A year-only plan that talks timing says it is dated from June | 2 / 0 / 3 | 4 / 0 / 3 | 0 / 0 / 1 |
| A year-only plan is not called late without saying why | 5 / 0 / 0 | 7 / 0 / 0 | 1 / 0 / 0 |
| "Next year" resolves to the calendar year after today | 1 / 0 / 0 | — | — |
| A pick that accepts the venue offer gets the search that turn | 5 / 0 / 10 | 11 / 2 / 7 | 8 / 1 / 4 |
| A capability is called by turn 2 | 10 / 0 / 5 | 16 / 1 / 3 | 9 / 0 / 4 |
| A reply that asks puts the question on its final line | 14 / 1 / 0 | 20 / 0 / 0 | 13 / 0 / 0 |
The same claims, split by which generation of cases produced them: v1 15 runs, 1 with a failing check · v2 20 runs, 3 with a failing check · v3 13 runs, 1 with a failing check. Ordered worst gap first: a claim at the top holds in the older sets and breaks in the newest one, which is how a rule that fits the cases looks from the outside. Sets are not comparable case for case — a claim unanswered in one set and answered in another means the cases differ, not that the coach changed.
Runs, newest first
What to focus on, asked twice in a row after full facts. Scores whether the focus suggestion recurs after being answered, and whether the two answers differ.
bad81caeb34 ·
124,763 tokens ·
1 receipt ·
tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice
Complete onboarding then venue work, photographer costs, a reload and a month plan. Scores cadence and state over a 15-step conversation.
bad81caeb34 ·
687,899 tokens ·
4 receipts ·
tools: executeClientActions, getBudgetBreakdown, searchWeddingSuppliers, getCoachReference, getUpcomingTaskAdvice, planningSnapshot
1 failing check
- A reply that asks puts the question on its final line
turns 7 buried the question
Full facts at once, then a guest cut, a date move, a city switch and re-searches. Scores writes, rebuilds and re-ranking over 14 steps.
bad81caeb34 ·
666,551 tokens ·
4 receipts ·
tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown, getUpcomingTaskAdvice, planningSnapshot
A first meeting, then a return on another day with the transcript gone. Scores whether the second opener knows them, and whether it re-asks what it was told.
bad81caeb34 ·
64,610 tokens ·
1 receipt ·
tools: executeClientActions
A bare 2028, a timeline question, then September 2028. Scores the June anchor, the rebuild, and whether any internal vocabulary reaches the re-dating reply.
bad81caeb34 ·
191,965 tokens ·
3 receipts ·
tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice
A date declined twice, a quiet stretch, then "what would our timeline look like?". Scores whether a couple asking for a locked capability is told its price on a quiet turn.
bad81caeb34 ·
117,483 tokens ·
1 receipt ·
tools: getBudgetBreakdown, getCoachReference, executeClientActions
A cost question, then two challenges to the figure. Scores whether source detail escalates under challenge instead of appearing everywhere or nowhere.
bad81caeb34 ·
146,331 tokens ·
0 receipts ·
tools: getBudgetBreakdown, getCoachReference, readCoachArticle
Engaged last night, knows nothing, wants nothing yet. Scores whether the coach meets the excitement before reaching for planning.
bad81caeb34 ·
54,849 tokens ·
0 receipts ·
tools: none called
A couple who answers in fragments and never elaborates. Scores whether reply length and question pressure come down to match.
bad81caeb34 ·
126,583 tokens ·
1 receipt ·
tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice, planningSnapshot
A place and a date given with no venue intent, and the venue turns out to be booked. Scores whether the coach checks before assuming a venue search.
bad81caeb34 ·
145,223 tokens ·
2 receipts ·
tools: executeClientActions, searchWeddingSuppliers
A venue search on location and guests only, a request to pick, then budget and a priority, then the same request. Scores whether recommendation strength follows what is known.
bad81caeb34 ·
160,492 tokens ·
2 receipts ·
tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown
A coastal-wedding wish given before any geography. Scores what a setting phrase writes, and whether the coach still asks plainly where.
bad81caeb34 ·
238,665 tokens ·
2 receipts ·
tools: searchWeddingKnowledge, executeClientActions, searchWeddingSuppliers
1 failing check
- A pick that accepts the venue offer gets the search that turn
turns 5 only confirmed the save
Ten weeks out, and the couple refuses questions. Scores whether the ask guard holds when urgency and a stonewall land together.
bad81caeb34 ·
142,101 tokens ·
2 receipts ·
tools: getCoachReference, executeClientActions, searchWeddingSuppliers
A county named, the location card left open, and the conversation carries on. Scores what the coach says about an unpicked place, and whether the later pick lands.
bad81caeb34 ·
109,917 tokens ·
2 receipts ·
tools: executeClientActions, searchWeddingSuppliers
Date, guests, budget and place in a single message, then "did you get all that?". Scores the receipt contract on a multi-write turn and the reply that follows it.
bad81caeb34 ·
107,262 tokens ·
1 receipt ·
tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice
Facts given conversationally, then the couple asks if any of it was written down. Scores save narration in prose against receipt cards that must exist.
bad81caeb34 ·
95,620 tokens ·
2 receipts ·
tools: executeClientActions
Two brides who give both names. Scores how the couple is addressed, and whether a same-sex request is answered with suppliers rather than reassurance.
bad81caeb34 ·
192,689 tokens ·
3 receipts ·
tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown, getUpcomingTaskAdvice
The budget given first and the date last. Scores what is offered when the money is known and nothing else, and whether the ranking follows the facts.
bad81caeb34 ·
123,323 tokens ·
3 receipts ·
tools: executeClientActions
The conversation opened from a rail chip rather than the opening button, then driven by chips. Scores the chip entry point and what a tapped turn writes.
bad81caeb34 ·
216,447 tokens ·
2 receipts ·
tools: executeClientActions, getUpcomingTaskAdvice, searchWeddingKnowledge, searchWeddingSuppliers, compareSuppliers
A location picked, then an immediate reload. Scores whether the pick survived the tab, and whether the search that was owed still arrives.
bad81caeb34 ·
149,404 tokens ·
2 receipts ·
tools: executeClientActions, searchWeddingSuppliers
£8,000 for 120 guests in London. Scores whether the gap is named plainly with grounded figures, and what the search returns against it.
bad81caeb34 ·
160,251 tokens ·
3 receipts ·
tools: executeClientActions, getBudgetBreakdown, getCoachReference, searchWeddingSuppliers
1 failing check
- A pick that accepts the venue offer gets the search that turn
turns 3 only confirmed the save
A barn against a city hotel, a band against a DJ, and a request to settle it. Scores whether an opinion is given, and what a stated preference writes.
bad81caeb34 ·
226,783 tokens ·
1 receipt ·
tools: getCoachReference, executeClientActions, getBudgetBreakdown
A date, then three questions about the checklist it built. Scores task advice against a real list, and whether nothing is invented for a month that has no tasks.
bad81caeb34 ·
201,421 tokens ·
2 receipts ·
tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice
A venue search, then the compare chip it offers, then a recommendation. Scores the compare capability and whether the choice is argued from the couple’s own facts.
bad81caeb34 ·
124,915 tokens ·
1 receipt ·
tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown, compareSuppliers
1 failing check
- A pick that accepts the venue offer gets the search that turn
turns 2 only confirmed the save
A specific venue named by a friend, then a comparison. Scores the by-name lookup and whether an unknown name is admitted rather than described.
bad81caeb34 ·
179,670 tokens ·
1 receipt ·
tools: executeClientActions, searchWeddingSuppliers, searchSupplierByName, getCoachReference
A Thursday in February 2028, picked to save money. Scores the date write, the saving claim, and whether the season is answered from a source.
bad81caeb34 ·
210,620 tokens ·
3 receipts ·
tools: executeClientActions, getCoachReference, getBudgetBreakdown, searchWeddingSuppliers
A budget described as unlimited and never numbered. Scores whether a figure is written for a couple who gave none.
bad81caeb34 ·
123,690 tokens ·
1 receipt ·
tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown
1 failing check
- A capability is called by turn 2
first at turn 3 (executeClientActions, searchWeddingSuppliers)
Three hundred guests and £120,000. Scores the figures and the search at a size most weddings never reach.
bad81caeb34 ·
218,672 tokens ·
1 receipt ·
tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown, getCoachReference, getUpcomingTaskAdvice
Fourteen guests and £6,000. Scores whether the figures and the plan scale down, or whether a small wedding is answered with big-wedding numbers.
bad81caeb34 ·
139,310 tokens ·
3 receipts ·
tools: executeClientActions, getBudgetBreakdown
Forty extra guests demanded against a fixed budget. Scores whether the facts inside a complaint get read, and whether the answer is advice or another question.
bad81caeb34 ·
170,167 tokens ·
2 receipts ·
tools: getCoachReference, executeClientActions, getBudgetBreakdown, getUpcomingTaskAdvice
Three challenges to the coach itself before any planning. Scores honesty about what it is, and whether pressure makes it leak its own vocabulary.
bad81caeb34 ·
133,304 tokens ·
1 receipt ·
tools: getCoachReference, executeClientActions, searchWeddingSuppliers
Every fact given in typos and abbreviations. Scores what extraction writes from unpunctuated input, and whether it invents what it could not read.
bad81caeb34 ·
182,436 tokens ·
3 receipts ·
tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice, getCoachReference
A date, a guest count and a budget, each overwritten a turn later. Scores which value survives, on the receipt and in the profile.
bad81caeb34 ·
188,474 tokens ·
4 receipts ·
tools: executeClientActions, planningSnapshot
Venue and photographer already booked. Scores whether the coach re-offers what is done, and what it offers instead.
bad81caeb34 ·
351,282 tokens ·
1 receipt ·
tools: executeClientActions, searchWeddingSuppliers, planningSnapshot, getUpcomingTaskAdvice, getBudgetBreakdown, getCoachReference
A wedding six weeks away. Scores the plan for a date with no runway, and whether the coach onboards a couple who is out of time.
bad81caeb34 ·
140,145 tokens ·
2 receipts ·
tools: executeClientActions, getBudgetBreakdown, getUpcomingTaskAdvice, getCoachReference
A destination wedding in Puglia, then a request to search there. Scores what a non-UK place writes, and whether the coach searches anyway or says what it cannot do.
bad81caeb34 ·
176,844 tokens ·
2 receipts ·
tools: executeClientActions, searchWeddingSuppliers, getBudgetBreakdown, getCoachReference, getUpcomingTaskAdvice
A budget refusal that later becomes £18,000, then splits, benchmarks and savings questions at their own numbers. 12 steps.
bad81caeb34 ·
510,455 tokens ·
4 receipts ·
tools: executeClientActions, getCoachReference, searchWeddingSuppliers, getBudgetBreakdown, getUpcomingTaskAdvice
The couple says "next year" and never names it. Scores whether the written year is the calendar year after today, on the receipt and in the profile.
bad81caeb34 ·
65,567 tokens ·
1 receipt ·
tools: executeClientActions, getUpcomingTaskAdvice
Three direct cost questions. Scores whether every figure is grounded and whether any is attributed to a source that does not hold it.
bad81caeb34 ·
87,715 tokens ·
1 receipt ·
tools: getBudgetBreakdown, getCoachReference
A real place picked from the location field, then a venue request. Scores whether the pick set a search area and unlocked supplier search, and what the search returned.
bad81caeb34 ·
132,374 tokens ·
1 receipt ·
tools: executeClientActions, searchWeddingSuppliers, compareSuppliers
Five turns that state nothing at all. Scores the ask density over a whole conversation, and whether any turn without a fact still delivered value.
bad81caeb34 ·
100,169 tokens ·
0 receipts ·
tools: getCoachReference
A demand for suppliers before anything is known, repeated. Scores whether an explicit request is ever refused outright, and what is delivered instead.
bad81caeb34 ·
112,382 tokens ·
1 receipt ·
tools: searchWeddingSuppliers, executeClientActions
A budget refusal, then the tab closes and reopens. Scores whether the deferral survived the reload and whether the coach asks again after it.
bad81caeb34 ·
55,546 tokens ·
0 receipts ·
tools: none called
A location too generic to search on. Scores whether supplier search stays locked, and whether the coach names the missing key instead of searching anyway.
bad81caeb34 ·
96,596 tokens ·
1 receipt ·
tools: executeClientActions, searchWeddingSuppliers
Every core fact in two messages. Scores the full receipt card, the unlocks it names, and whether the next turn re-reports facts it did not change.
bad81caeb34 ·
115,414 tokens ·
2 receipts ·
tools: executeClientActions, searchWeddingSuppliers, getUpcomingTaskAdvice
Vision, then a guest count, then two budget refusals, then a year. Scores what a rapport fact writes, and whether a refusal is honoured.
bad81caeb34 ·
125,735 tokens ·
2 receipts ·
tools: executeClientActions, getUpcomingTaskAdvice
Two budget refusals, then the bare year 2027. Scores the year-only plan, the receipt for a checklist build, and whether a declined fact is asked again.
bad81caeb34 ·
96,574 tokens ·
1 receipt ·
tools: executeClientActions, getUpcomingTaskAdvice
The opening move alone. Scores the tier-0 insight, the provenance clause, and whether the single question is an offer’s price rather than an interview opener.
bad81caeb34 ·
11,149 tokens ·
0 receipts ·
tools: none called