Kopi Uncle Β· Ten-Tester Validation

First-Play Docket

A day-by-day working copy of the Ten-Tester WhatsApp Validation Proposal β€” every setup step, WhatsApp script, and form/sheet spec ready to use in order. Checkboxes and the roster save in this browser as you go.

πŸ‘₯ 10 valid sessions
πŸ” β‰₯6 unprompted retries
⏱ 30–90s median first run
🚫 0 open S1 issues
BEFORE DAY 1

What this test decides

One decision only: is the current game loop strong enough for a soft launch. Not a marketing launch, not a feature vote. Full policy lives in docs/validation/TEN_TESTER_WHATSAPP_VALIDATION_PROPOSAL.md β€” this page is the working copy for actually running it.

The four-part gate

All four must hold, or the result is Redesign / Invalid β€” never "close enough."

Why WhatsApp, and the one hard rule

One temporary group handles recruiting, scheduling, and the debrief β€” familiar to the Singapore audience, low-friction to join. The rule that protects the data: only the moderator posts until every form is locked. Testers submit privately; an early "this was easy!" in the group teaches later testers how to play and pressures them toward the same answer. That single leak is the most common way this kind of test quietly breaks.

Decision needed before Day 1 β€” how is timing/retry measured?

Self-report (default, zero setup): testers estimate their own first-run seconds and run count on the form. The proposal explicitly allows this β€” it's just the lowest-confidence tier of evidence. Fine for a 10-person friends-and-network test if you stay honest about that lower confidence in the final report.

Lightweight telemetry (higher confidence, ~30–60 min of engineering): I can add a few anonymous timestamped events (game start, first Game Over, retry) so the median and retry count come from real clocks, not memory. No new tracking identifier β€” it would reuse the same anonymous client ID Analytics.js already sends.

Recommendation: for twelve known, friendly testers, self-report plus asking two or three of them to screen-record is proportionate. Say the word if you'd rather I build the telemetry first β€” it's a short job and can still land before Day 3.

The week at a glance

DayPhaseWhat happensOwner effort
1SetupForm, Sheet, WhatsApp group, production smoke test~3h
2SetupOptional telemetry, final freeze on production1–4h
3RecruitTen primaries + two reserves confirmed and ID'd~1h
4–5RunStaggered independent sessions2–3h total
6Close outReconcile, calculate the gate, open debrief~2h
7Close outPublish the decision, update the project record~1h
DAY 1–2

Build the Form, Sheet, and group

Nothing here touches a tester yet β€” this is the scaffolding the whole test writes into.

Setup checklist

Google Form β€” "Kopi Uncle First-Play Test"

Behavioral questions before open-ended ones, so later questions can't reshape earlier behavior. Never ask "did you want to retry?" before recording the run count.

Section A β€” Consent & eligibility

1Test IDrequired Β· pattern KU-T01–KU-T12
2I am at least 18 and agree to this short product testrequired yes/no
3I understand participation is voluntary and I may stoprequired yes/no
4Have you previously played or watched Kopi Uncle?required yes/no
5Evidence methodevents / recording / observed / self-report only
6DeviceiPhone / Android / other
7BrowserSafari / Chrome / other
8Kopi shorthand familiarity before todaylow / medium / high

Section B β€” Uncoached first play

1Did the game load?yes/no
2Did you reach Game Over?yes/no
3How many runs did you start before opening this form?required number
4Approximate first-run durationrequired seconds if no telemetry
5First-run scoreoptional number
6Which order or control first caused confusion?short answer
7What did you think the objective was?short answer

Section C β€” End-to-end checks (each: Worked / Failed / Did not attempt + optional error note)

1Submit score and display name
2Open Hall of Fame
3Find the submitted score or understand its result
4Generate share card
5Open the native share route or download fallback
6See the email signup prompt
7Submit email only if voluntarily chosen

Section D β€” Feedback

1How clear was the first order?1–5
2How satisfying did a correct order feel?1–5
3How fair did a wrong order feel?1–5
4How likely are you to play again tomorrow?0–10 Β· diagnostic only, not a gate
5What was the single best part?short answer
6What was the single most confusing/frustrating part?short answer
7If one thing changed before launch, what should it be?short answer
8May an anonymized quotation be used in the report?yes/no

Google Sheet β€” columns & gate formulas

One row per tester, rows 2–13 (ten primaries + two reserves).

A–ETest ID Β· Valid record (TRUE/FALSE) Β· Exclusion reason Β· Device/browser Β· Familiarity
F–IUnprompted retry (TRUE/FALSE) Β· First-run seconds Β· Run count Β· First-run score
J–MScore submission outcome Β· Hall of Fame outcome Β· Share outcome Β· Signup-flow outcome
N–OHighest issue severity Β· Moderator notes
Docket Β· Sheet formulas
Valid sessions
=COUNTIF(B2:B13,TRUE)

Verified unprompted retries
=COUNTIFS(B2:B13,TRUE,F2:F13,TRUE)

Median valid first-run seconds
=MEDIAN(FILTER(G2:G13,B2:B13=TRUE))

Behavioral gate
=AND(COUNTIF(B2:B13,TRUE)=10,
     COUNTIFS(B2:B13,TRUE,F2:F13,TRUE)>=6,
     MEDIAN(FILTER(G2:G13,B2:B13=TRUE))>=30,
     MEDIAN(FILTER(G2:G13,B2:B13=TRUE))<=90)

Keep exclusions visible β€” never delete an inconvenient row or quietly swap in a reserve without marking the original FALSE.

WhatsApp group settings

Pre-test production smoke test

The full deep-test suite already passed on 9 Sep 2026 β€” this is a quick re-check the moment before your first tester, since anything could have shifted since.

DAY 3

Find ten primaries, two reserves

Recruit privately by WhatsApp DM β€” no public link, no group forwarding. The group itself forms only after people have already agreed to take part.

Target mix

DimensionTarget
Mobile platformAt least 3 iPhones and 3 Android phones
Kopi shorthand familiarityAt least 3 low, 3 medium, 2 high
RelationshipNo more than 5 close friends/family β€” include colleagues or weak-tie contacts
AgeMix of younger and older adult mobile users where practical

Never swap out a tester for playing poorly β€” only for withdrawing, prior exposure, coaching, or a confirmed technical failure.

Where the weak-tie half comes from

Your own chats skew toward people who already know you β€” and often already know some kopi shorthand. To hit the familiarity spread without posting publicly, try: a colleague WhatsApp group, a condo/CC or alumni group chat, a parent-group chat, or a neighbour you'd DM directly. Each still goes through the same private-DM screening below β€” the difference is just whose network you're borrowing.

Eligibility, screened by DM before anyone joins the group

Recruitment message β€” send by DM, one at a time

Docket Β· DM
I'm testing a short Singapore kopi-ordering mobile game. I need about 15 minutes of first-play feedback from people who have not seen it before. There is no need to be good at games or kopi shorthand. You will play on your own phone without coaching, try a few end-to-end actions, then complete a private form. Joining is voluntary, and the group will be used only to coordinate this test. Please do not forward the game link. Reply privately if you can take part.

Roster tracker

Saved only in this browser β€” nothing here leaves your machine.

Test IDDeviceFamiliarityStatusNotes
DAY 4–5

Run the sessions

Start testers in staggered windows. Send only the neutral scripts below β€” no hints, no mention of retry, timing, or kopi terms.

Welcome message β€” post once, pin it

Docket Β· Group
Thank you for helping test Kopi Uncle. Please wait for your assigned start message. During the test, play independently and do not ask the group for hints. Do whatever feels natural after the first Game Over. Then complete the private form. Do not post your score or opinions here until I reopen discussion. If the page will not load, message me privately with a screenshot.

Individual start message β€” DM each tester their own copy

Docket Β· DM Β· fill in KU-T__ and the form link
Your test ID is KU-T__. Turn on Do Not Disturb if you are recording your screen. Open this link and use the game as you naturally would: https://play.3bearskopi.com. When you feel finished, complete this private form: [FORM LINK]. I will not answer gameplay questions during the test, but message me privately if the site cannot load.

Never mention retry, session length, target scores, or what any kopi term means in this message.

Submission reminder β€” post as the deadline nears

Docket Β· Group Β· fill in the time
A quick reminder to submit the private form by [TIME] if you completed your test. Please still avoid discussing the game in this group until submissions are locked.

Run-day checklist

Severity & stop rules

SevDefinitionResponse
S1Can't load/start/finish, score corruption, privacy/security exposure, widespread data lossStop testing, preserve evidence, fix, restart the affected run
S2A major path fails but the tester can continue (e.g. leaderboard, share)Log immediately; gate can still pass, but soft launch is blocked until resolved
S3Confusing copy, isolated visual defect, slow response, recoverable errorLog and prioritize after the behavioral result
S4Preference or enhancement requestRecord without changing the test build

Pause the whole test if more than one tester hits the same S1/S2, the deployed revision changes, or a group message leaks gameplay guidance before later testers finish.

DAY 6–7

Reconcile, decide, publish

Calculate the gate before you read the group's discussion β€” a good story afterward shouldn't move a number decided beforehand.

Debrief message β€” post only after every valid form is locked

Docket Β· Group
All valid responses are now lockedβ€”thank you. Discussion is open. What surprised you, what was confusing, and what made you continue or stop? Please avoid posting phone numbers, email addresses, or screenshots containing notifications.

Close-out checklist

Decision rules

ResultConditionsNext action
Go10 valid sessions Β· retries β‰₯6 Β· median 30–90s Β· no open S1Resolve S2s, then a controlled warm-channel soft launch
Conditional goGate passes but one narrow, fixable defect showed upFix that defect only, run a focused confirmation, then soft launch
RedesignRetries <6 or median outside 30–90sDiagnose the failure point, redesign the loop, retest with ten new testers
Invalid β€” retestFewer than 10 valid, coaching contamination, a mid-test deploy, or shaky timing dataFix the study flaw, replace only the invalid sessions

Never average away a failed threshold, and never let survey enthusiasm substitute for an observed retry.

Deliverables package

Bring me the results when you have them β€” I'll write up the decision report and update PROJECT.md's decision log and action register with you.