NataCoach / product & system design Wiki Brain Personas Coach Console ↗

11 · Roadmap

Four phases from a hand-driven concierge pilot to 100+ clients, each gated by measurable exit criteria — dates are planning assumptions, gates are not.


0. How to read this roadmap

Three rules govern the plan:

  1. Phases exit on criteria, not calendar. The dates below assume a 2026-09-01 start and are for coordination; a phase that misses its exit criteria does not proceed, it iterates or kills the failing assumption (see §6).
  2. Autonomy is earned, not scheduled. The coach-in-the-loop dial (draft → auto-send with audit → fully autonomous, per message type) only moves when the measured edit rate for that message type clears its bar. The phases below say when we expect the dial to move; the Coach Console edit-rate data says whether it does. Safety-critical messages stay at escalate in every phase, forever.
  3. Build the relationship before the machinery. Phase 0 deliberately ships almost no automation. If 10 humans will not photograph meals and tap ✓ for a hand-driven version of the product, no pipeline will save it.

A note on dates: the worked examples across docs 0210 (Marta's week 8, Denys's Lisbon week, dated Jul–Aug 2026) depict the Phase-2 steady state — every feature live, autonomy dials graduated — rendered on 2026 calendar dates for readability. They are illustrations of what the product feels like once Phase 2 exits, not snapshots of this build calendar.

Phase Name Window (assumption) Clients Autonomy dial Headline
0 Concierge pilot Sep–Oct 2026 (2 mo) 10 Everything draft Bot relays, Nata approves every word
1 Automation core Nov 2026–Jan 2027 (3 mo) 10 Routine types auto-send with audit Meal Lens, Session Mode, Health Sync live
2 Rich experiences Feb–Apr 2027 (3 mo) 10 25 Routine types autonomous Mini App runner, Form Check, Weekly Review
3 Scale to 100+ May–Aug 2027 (4 mo) 25 100+ Mostly open Squads, payments, ops hardening

1. Phase 0 — "Concierge pilot" (Sep–Oct 2026)

Thesis being tested: clients want this relationship in this channel, and Nata's judgment can be captured as edits on drafts.

Scope

  • @NataCoachBot live: onboarding conversation (goals, injuries, contraindications, schedule), photo intake, buttons, voice notes. Chat-first only — no Mini App yet.
  • Bot relays; Nata approves everything. Every outbound coaching message — meal feedback, workout adjustments, check-in replies — is an LLM draft in Nata's approval queue (Coach Console v0 = queue + per-client timeline, nothing else). Dial fully at draft.
  • Meal photos flow through a manual-assist Meal Lens: LLM proposes macros + feedback draft; Nata approves/edits; median item time < 30 s is the design target.
  • Workouts delivered as structured chat messages from Nata's programs (authored in a shared sheet for now); completion + RPE captured with buttons. No proposal engine yet — weights come straight from the program.
  • No Health Sync integration yet: recovery captured conversationally or via user screenshots of Whoop/Garmin (deliberately painful — it sizes the value of Phase 1).
  • Event log in Postgres from day one (locked decision 6) + nightly Wiki Brain distillation v0 (raw/, wiki/, index.md, log.md per the Karpathy method). Memory compounds from the first message even while automation is absent.
  • Safety keyword escalation live from day one.

Exit criteria (measured over pilot weeks 5–8)

Criterion Bar
Clients onboarded and active (≥ 4 interactions/week) 10 / 10 onboarded, ≥ 8 active
Meal-photo behavior exists without automation ≥ 60% of days with ≥ 2 photos
Draft quality baseline established Edit-rate dashboard live; ≥ 60% of drafts approved unedited
Approval queue is fast enough to be sustainable Median approval < 30 s/item; Nata total ≤ 60 min/client/week (down from ~110)
Coaching stays in-channel ≥ 90% of client interactions via @NataCoachBot, not Nata's personal DMs
Safety rail fires correctly 100% of seeded + real pain/injury phrases escalated

2. Phase 1 — "Automation core" (Nov 2026–Jan 2027)

Thesis being tested: automation raises adherence and quality while cutting Nata's time ~4×.

Scope

  • Health Sync live: Whoop API v2 + Garmin Health API, OAuth deep link from chat, webhook ingestion. Screenshots retired.
  • Meal Lens live end-to-end: photo macros coach-grade feedback vs. Nata's food program, propose-confirm on macros, < 60 s photo-to-feedback.
  • Session Mode v1 in chat: guided exercise-by-exercise runner with buttons; weight-proposal engine (history + today's recovery, progressive overload); RPE capture per exercise.
  • Morning Brief live: recovery-driven daily plan proposal before 07:30 local.
  • Graduated autonomy begins: message types move to auto-send with audit via the per-client streak mechanics in 09 §3.3 (20 consecutive unedited approvals per client, promotion confirmed by Nata); fleet-level we expect ≥ 90% approve-unedited over a rolling 100 items before most clients hit their streaks (expected first: meal feedback on green-recovery days, session confirmations, routine nudges). Program changes and anything safety-adjacent stay at draft/escalate.
  • Coach Console v1: roster health view, alert feed, autonomy dial controls, cost telemetry.

Exit criteria — this phase owns the pilot success metrics:

Criterion Bar
Training adherence ≥ 80% of planned sessions
Meal-log coverage ≥ 70% of days (≥ 2 meals via Meal Lens)
Nata time per client ≤ 30 min/week, instrumented
Retention ≥ 8 / 10 opt to continue past week 8 of this phase
Proposal acceptance (weights, macros, adjustments) ≥ 70% confirmed unedited
≥ 2 message types running at auto-send with audit With < 5% post-hoc correction rate
Health Sync freshness Recovery data present for ≥ 90% of Morning Briefs by 07:30
Unit cost LLM spend < $2/user/day

3. Phase 2 — "Rich experiences" (Feb–Apr 2027)

Thesis being tested: rich surfaces deepen the relationship (not just decorate it), and Form Check works without custom CV models.

Scope

  • Mini App workout runner: Next.js Telegram Mini App — live set/rest timers, plate math, one-thumb logging; chat runner remains as fallback (chat-first principle).
  • Form Check live: video upload vision-LLM technique analysis against per-exercise checklists 2–3 coaching cues, Nata-reviewed at draft until it earns its own dial movement; hard cases escalate with timestamped clips.
  • Weekly Review: Sunday chart-rich summary for the client (Mini App) + cross-client digest for Nata.
  • Coach Console v2: full analytics, program builder (replacing the sheet), chat takeover.
  • Apple Health via bridge (e.g. Health Auto Export) for non-Whoop/Garmin clients.
  • First-tier autonomy: message types move to fully autonomous via the per-client mechanics in 09 §3.3 (30 audited sends, 0 retractions, ≥ 14 days in audit); the fleet-level expectation is < 3% correction over 300 audited items before promotions become routine. Intake grows 10 25 clients (waitlist) to stress the queue before Phase 3.

Exit criteria

Criterion Bar
Mini App runner adoption ≥ 60% of sessions run in Mini App (rest in chat) — with session completion no worse than chat-only
Form Check usefulness Nata agrees with the top cue on ≥ 75% of reviewed videos; ≥ 1 video/client/month uploaded
Weekly Review engagement ≥ 70% of clients open it within 24 h
Autonomy ≥ 2 message types fully autonomous with < 3% correction rate
Scale rehearsal 25 active clients; Nata time ≤ 20 min/client/week; adherence and coverage bars from Phase 1 still met
Perceived personalness ≥ 8/10 clients answer "feels like Nata is coaching me" ≥ 4/5 in monthly survey

4. Phase 3 — "Scale to 100+" (May–Aug 2027)

Thesis being tested: the economics — one coach, 100+ clients, people paying, relationship intact.

Scope

  • Autonomy dial mostly open: everyday coaching (meal feedback, session coaching, briefs, reviews) fully autonomous with sampled audits (Nata reviews a random 5%); draft reserved for program changes and plateau calls; safety always escalate.
  • Squads: optional small group chats (5–8 clients, e.g. "postpartum strength", "fat-loss travelers") for community and challenges — the optional later from locked decision 1; 1:1 private chat remains the coaching channel.
  • Payments: subscription billing at the target $79–129/month (provider chosen then — Stripe or Telegram-native; deliberately out of scope until now per brief §6).
  • Ops hardening: rate-limit-aware queues, per-user cost caps, on-call alerting, data-export/delete self-service, Console triage views built for a 100-roster (exception-based, not per-client).
  • Growth 25 100+ via waitlist and squad referrals.

Exit criteria

Criterion Bar
Roster ≥ 100 active paying clients
Nata time ≤ 10 min/client/week at the 100-client mark
Retention ≥ 85% month-over-month across the paid base
Adherence / coverage Phase 1 bars (80% / 70%) hold at 100+ clients
Economics Gross margin per client positive with LLM < $2/user/day; payment conversion from pilot cohort ≥ 70%
Trust under autonomy Complaint/correction rate on autonomous messages < 2%; safety SLA still 100%

5. Timeline

gantt
 title NataCoach roadmap (planning assumption, start 2026-09-01)
 dateFormat YYYY-MM-DD
 axisFormat %b %Y
 section Phase 0 Concierge pilot
 Bot onboarding and approval queue :p0a, 2026-09-01, 21d
 Seed Wiki Brains and programs :p0b, 2026-09-08, 14d
 Run 10-client concierge pilot :p0c, 2026-09-22, 34d
 Gate review and edit-rate baseline :p0d, 2026-10-26, 6d
 section Phase 1 Automation core
 Health Sync Whoop and Garmin :p1a, 2026-11-01, 35d
 Meal Lens end to end :p1b, 2026-11-01, 42d
 Session Mode chat runner :p1c, 2026-11-17, 42d
 Morning Brief and proposal engine :p1d, 2026-12-08, 28d
 Graduated autonomy rollout :p1e, 2027-01-05, 26d
 section Phase 2 Rich experiences
 Mini App workout runner :p2a, 2027-02-01, 42d
 Form Check pipeline :p2b, 2027-02-15, 42d
 Coach Console full analytics :p2c, 2027-03-01, 45d
 Weekly Review and 25-client intake :p2d, 2027-03-16, 45d
 section Phase 3 Scale to 100 plus
 Autonomy mostly open :p3a, 2027-05-01, 45d
 Squads and community :p3b, 2027-06-01, 45d
 Payments and billing :p3c, 2027-06-15, 45d
 Grow roster to 100 plus :p3d, 2027-05-01, 122d

6. Riskiest assumptions, and how each phase tests them

Ordered by how fatal being wrong would be. "Kill/pivot signal" is the number at which we stop pushing and change the design instead.

# Assumption If it is false How the phases test it Kill / pivot signal
1 Clients will photograph meals and tap ✓ consistently when friction is near zero Meal Lens and the food program are decorative; half the product's data loop dies P0 measures raw willingness with a hand-driven pipeline (bar: 60% of days); P1 must lift it to 70% with instant feedback; P3 must hold it at 100 clients Coverage < 50% in P0 despite < 15 s per meal pivot to voice/text meal capture or drop per-meal scoring for daily summaries
2 An LLM with the Wiki Brain + Nata's programs can draft feedback Nata sends mostly unedited Coach-in-the-loop never graduates; Nata stays the bottleneck; economics collapse P0 establishes the edit-rate baseline (bar: 60% unedited); P1 requires 90% for dial movement; P2/P3 require sustained < 3% / < 2% correction under autonomy Unedited rate stuck < 75% after P1 prompt/persona iteration keep humans in loop permanently and re-price, or narrow autonomous scope to confirmations only
3 Nata's time can compress ~110 30 10 min/client/week without clients feeling the difference The product is a nicer workflow tool, not a scaling engine P0 instruments her time (bar: ≤ 60); P1 gates on ≤ 30; P2 adds the perceived-personalness survey (≥ 4/5); P3 gates on ≤ 10 at 100 clients Time plateaus > 40 min/client or personalness < 3.5/5 cap roster ~30 and reposition as premium hybrid coaching
4 Recovery-aware proactive coaching improves adherence (vs. annoying people) Morning Brief is spam; wearable integration is a gimmick P0 ships no integration (baseline adherence + screenshot pain); P1 compares adherence and Brief response rates before/after Health Sync; mute-rate tracked Brief response < 40% or mute rate > 20% in P1 reduce to 2–3 proactive moments/week, triggered only on red-flag recovery
5 Vision LLM + per-exercise checklists give useful form feedback without custom pose models Form Check ships embarrassing cues or nothing P2 gates on Nata agreeing with the top cue ≥ 75%; disagreements become checklist edits, logged weekly Agreement < 60% after two checklist iterations Form Check becomes "video routed to Nata with AI pre-notes", revisit pose models post-P3
6 Unit economics work: LLM < $2/user/day and clients pay $79–129/month Scaling multiplies losses P1 instruments per-user cost with model routing (Opus-tier reasoning, Haiku-tier routing); P3 tests willingness to pay (conversion ≥ 70% from pilot cohort) Cost > $3/user/day after routing/caching work, or conversion < 50% cut always-on surfaces, raise price, or gate expensive pipelines to higher tiers
7 Trust survives autonomy — clients accept that everyday messages are AI, backed by Nata Dial-opening triggers churn exactly when economics need it P1 audits every auto-sent message; P2 survey after first autonomous types; P3 watches complaint rate (< 2%) and retention through full opening Retention dips > 10 pts within a month of a dial change roll the dial back one notch and re-earn with a longer audit window
8 Whoop/Garmin APIs are reliable enough for a 07:30 Morning Brief The flagship proactive moment misfires on stale data P1 gates on 90% data freshness; degraded mode ("no sync yet — how did you sleep?") specified in Health Sync Freshness < 80% sustained decouple Brief from sync (send on schedule, enrich when data lands)

The vision this plan serves is in 01 · Product Vision; what each shipped phase feels like for Marta and Denys is in 02 · User Experience and 10 · Personas.