A day-by-day working copy of the Ten-Tester WhatsApp Validation Proposal β every setup step, WhatsApp script, and form/sheet spec ready to use in order. Checkboxes and the roster save in this browser as you go.
One decision only: is the current game loop strong enough for a soft launch. Not a marketing
launch, not a feature vote. Full policy lives in
docs/validation/TEN_TESTER_WHATSAPP_VALIDATION_PROPOSAL.md β this page is the
working copy for actually running it.
All four must hold, or the result is Redesign / Invalid β never "close enough."
One temporary group handles recruiting, scheduling, and the debrief β familiar to the Singapore audience, low-friction to join. The rule that protects the data: only the moderator posts until every form is locked. Testers submit privately; an early "this was easy!" in the group teaches later testers how to play and pressures them toward the same answer. That single leak is the most common way this kind of test quietly breaks.
Self-report (default, zero setup): testers estimate their own first-run seconds and run count on the form. The proposal explicitly allows this β it's just the lowest-confidence tier of evidence. Fine for a 10-person friends-and-network test if you stay honest about that lower confidence in the final report.
Lightweight telemetry (higher confidence, ~30β60 min of engineering): I can
add a few anonymous timestamped events (game start, first Game Over, retry) so the median
and retry count come from real clocks, not memory. No new tracking identifier β it would
reuse the same anonymous client ID Analytics.js already sends.
Recommendation: for twelve known, friendly testers, self-report plus asking two or three of them to screen-record is proportionate. Say the word if you'd rather I build the telemetry first β it's a short job and can still land before Day 3.
| Day | Phase | What happens | Owner effort |
|---|---|---|---|
| 1 | Setup | Form, Sheet, WhatsApp group, production smoke test | ~3h |
| 2 | Setup | Optional telemetry, final freeze on production | 1β4h |
| 3 | Recruit | Ten primaries + two reserves confirmed and ID'd | ~1h |
| 4β5 | Run | Staggered independent sessions | 2β3h total |
| 6 | Close out | Reconcile, calculate the gate, open debrief | ~2h |
| 7 | Close out | Publish the decision, update the project record | ~1h |
Nothing here touches a tester yet β this is the scaffolding the whole test writes into.
Behavioral questions before open-ended ones, so later questions can't reshape earlier behavior. Never ask "did you want to retry?" before recording the run count.
Section A β Consent & eligibility
| 1 | Test ID | required Β· pattern KU-T01βKU-T12 |
| 2 | I am at least 18 and agree to this short product test | required yes/no |
| 3 | I understand participation is voluntary and I may stop | required yes/no |
| 4 | Have you previously played or watched Kopi Uncle? | required yes/no |
| 5 | Evidence method | events / recording / observed / self-report only |
| 6 | Device | iPhone / Android / other |
| 7 | Browser | Safari / Chrome / other |
| 8 | Kopi shorthand familiarity before today | low / medium / high |
Section B β Uncoached first play
| 1 | Did the game load? | yes/no |
| 2 | Did you reach Game Over? | yes/no |
| 3 | How many runs did you start before opening this form? | required number |
| 4 | Approximate first-run duration | required seconds if no telemetry |
| 5 | First-run score | optional number |
| 6 | Which order or control first caused confusion? | short answer |
| 7 | What did you think the objective was? | short answer |
Section C β End-to-end checks (each: Worked / Failed / Did not attempt + optional error note)
| 1 | Submit score and display name |
| 2 | Open Hall of Fame |
| 3 | Find the submitted score or understand its result |
| 4 | Generate share card |
| 5 | Open the native share route or download fallback |
| 6 | See the email signup prompt |
| 7 | Submit email only if voluntarily chosen |
Section D β Feedback
| 1 | How clear was the first order? | 1β5 |
| 2 | How satisfying did a correct order feel? | 1β5 |
| 3 | How fair did a wrong order feel? | 1β5 |
| 4 | How likely are you to play again tomorrow? | 0β10 Β· diagnostic only, not a gate |
| 5 | What was the single best part? | short answer |
| 6 | What was the single most confusing/frustrating part? | short answer |
| 7 | If one thing changed before launch, what should it be? | short answer |
| 8 | May an anonymized quotation be used in the report? | yes/no |
One row per tester, rows 2β13 (ten primaries + two reserves).
| AβE | Test ID Β· Valid record (TRUE/FALSE) Β· Exclusion reason Β· Device/browser Β· Familiarity |
| FβI | Unprompted retry (TRUE/FALSE) Β· First-run seconds Β· Run count Β· First-run score |
| JβM | Score submission outcome Β· Hall of Fame outcome Β· Share outcome Β· Signup-flow outcome |
| NβO | Highest issue severity Β· Moderator notes |
Valid sessions
=COUNTIF(B2:B13,TRUE)
Verified unprompted retries
=COUNTIFS(B2:B13,TRUE,F2:F13,TRUE)
Median valid first-run seconds
=MEDIAN(FILTER(G2:G13,B2:B13=TRUE))
Behavioral gate
=AND(COUNTIF(B2:B13,TRUE)=10,
COUNTIFS(B2:B13,TRUE,F2:F13,TRUE)>=6,
MEDIAN(FILTER(G2:G13,B2:B13=TRUE))>=30,
MEDIAN(FILTER(G2:G13,B2:B13=TRUE))<=90)
Keep exclusions visible β never delete an inconvenient row or quietly swap in a reserve without marking the original FALSE.
The full deep-test suite already passed on 9 Sep 2026 β this is a quick re-check the moment before your first tester, since anything could have shifted since.
Recruit privately by WhatsApp DM β no public link, no group forwarding. The group itself forms only after people have already agreed to take part.
| Dimension | Target |
|---|---|
| Mobile platform | At least 3 iPhones and 3 Android phones |
| Kopi shorthand familiarity | At least 3 low, 3 medium, 2 high |
| Relationship | No more than 5 close friends/family β include colleagues or weak-tie contacts |
| Age | Mix of younger and older adult mobile users where practical |
Never swap out a tester for playing poorly β only for withdrawing, prior exposure, coaching, or a confirmed technical failure.
Your own chats skew toward people who already know you β and often already know some kopi shorthand. To hit the familiarity spread without posting publicly, try: a colleague WhatsApp group, a condo/CC or alumni group chat, a parent-group chat, or a neighbour you'd DM directly. Each still goes through the same private-DM screening below β the difference is just whose network you're borrowing.
I'm testing a short Singapore kopi-ordering mobile game. I need about 15 minutes of first-play feedback from people who have not seen it before. There is no need to be good at games or kopi shorthand. You will play on your own phone without coaching, try a few end-to-end actions, then complete a private form. Joining is voluntary, and the group will be used only to coordinate this test. Please do not forward the game link. Reply privately if you can take part.
Saved only in this browser β nothing here leaves your machine.
| Test ID | Device | Familiarity | Status | Notes |
|---|
Start testers in staggered windows. Send only the neutral scripts below β no hints, no mention of retry, timing, or kopi terms.
Thank you for helping test Kopi Uncle. Please wait for your assigned start message. During the test, play independently and do not ask the group for hints. Do whatever feels natural after the first Game Over. Then complete the private form. Do not post your score or opinions here until I reopen discussion. If the page will not load, message me privately with a screenshot.
Your test ID is KU-T__. Turn on Do Not Disturb if you are recording your screen. Open this link and use the game as you naturally would: https://play.3bearskopi.com. When you feel finished, complete this private form: [FORM LINK]. I will not answer gameplay questions during the test, but message me privately if the site cannot load.
Never mention retry, session length, target scores, or what any kopi term means in this message.
A quick reminder to submit the private form by [TIME] if you completed your test. Please still avoid discussing the game in this group until submissions are locked.
| Sev | Definition | Response |
|---|---|---|
| S1 | Can't load/start/finish, score corruption, privacy/security exposure, widespread data loss | Stop testing, preserve evidence, fix, restart the affected run |
| S2 | A major path fails but the tester can continue (e.g. leaderboard, share) | Log immediately; gate can still pass, but soft launch is blocked until resolved |
| S3 | Confusing copy, isolated visual defect, slow response, recoverable error | Log and prioritize after the behavioral result |
| S4 | Preference or enhancement request | Record without changing the test build |
Pause the whole test if more than one tester hits the same S1/S2, the deployed revision changes, or a group message leaks gameplay guidance before later testers finish.
Calculate the gate before you read the group's discussion β a good story afterward shouldn't move a number decided beforehand.
All valid responses are now lockedβthank you. Discussion is open. What surprised you, what was confusing, and what made you continue or stop? Please avoid posting phone numbers, email addresses, or screenshots containing notifications.
| Result | Conditions | Next action |
|---|---|---|
| Go | 10 valid sessions Β· retries β₯6 Β· median 30β90s Β· no open S1 | Resolve S2s, then a controlled warm-channel soft launch |
| Conditional go | Gate passes but one narrow, fixable defect showed up | Fix that defect only, run a focused confirmation, then soft launch |
| Redesign | Retries <6 or median outside 30β90s | Diagnose the failure point, redesign the loop, retest with ten new testers |
| Invalid β retest | Fewer than 10 valid, coaching contamination, a mid-test deploy, or shaky timing data | Fix the study flaw, replace only the invalid sessions |
Never average away a failed threshold, and never let survey enthusiasm substitute for an observed retry.
Bring me the results when you have them β I'll write up the decision report and update PROJECT.md's decision log and action register with you.