0. How to read this roadmap
Three rules govern the plan:
- Phases exit on criteria, not calendar. The dates below assume a 2026-09-01 start and are for coordination; a phase that misses its exit criteria does not proceed, it iterates or kills the failing assumption (see §6).
- Autonomy is earned, not scheduled. The coach-in-the-loop dial (
draft → auto-send with audit → fully autonomous, per message type) only moves when the measured edit rate for that message type clears its bar. The phases below say when we expect the dial to move; the Coach Console edit-rate data says whether it does. Safety-critical messages stay atescalatein every phase, forever. - Build the relationship before the machinery. Phase 0 deliberately ships almost no automation. If 10 humans will not photograph meals and tap ✓ for a hand-driven version of the product, no pipeline will save it.
A note on dates: the worked examples across docs 02–10 (Marta's week 8, Denys's Lisbon week, dated Jul–Aug 2026) depict the Phase-2 steady state — every feature live, autonomy dials graduated — rendered on 2026 calendar dates for readability. They are illustrations of what the product feels like once Phase 2 exits, not snapshots of this build calendar.
| Phase | Name | Window (assumption) | Clients | Autonomy dial | Headline |
|---|---|---|---|---|---|
| 0 | Concierge pilot | Sep–Oct 2026 (2 mo) | 10 | Everything draft |
Bot relays, Nata approves every word |
| 1 | Automation core | Nov 2026–Jan 2027 (3 mo) | 10 | Routine types → auto-send with audit |
Meal Lens, Session Mode, Health Sync live |
| 2 | Rich experiences | Feb–Apr 2027 (3 mo) | 10 → 25 | Routine types → autonomous |
Mini App runner, Form Check, Weekly Review |
| 3 | Scale to 100+ | May–Aug 2027 (4 mo) | 25 → 100+ | Mostly open | Squads, payments, ops hardening |
1. Phase 0 — "Concierge pilot" (Sep–Oct 2026)
Thesis being tested: clients want this relationship in this channel, and Nata's judgment can be captured as edits on drafts.
Scope
- @NataCoachBot live: onboarding conversation (goals, injuries, contraindications, schedule), photo intake, buttons, voice notes. Chat-first only — no Mini App yet.
- Bot relays; Nata approves everything. Every outbound coaching message — meal feedback, workout adjustments, check-in replies — is an LLM draft in Nata's approval queue (Coach Console v0 = queue + per-client timeline, nothing else). Dial fully at
draft. - Meal photos flow through a manual-assist Meal Lens: LLM proposes macros + feedback draft; Nata approves/edits; median item time < 30 s is the design target.
- Workouts delivered as structured chat messages from Nata's programs (authored in a shared sheet for now); completion + RPE captured with buttons. No proposal engine yet — weights come straight from the program.
- No Health Sync integration yet: recovery captured conversationally or via user screenshots of Whoop/Garmin (deliberately painful — it sizes the value of Phase 1).
- Event log in Postgres from day one (locked decision 6) + nightly Wiki Brain distillation v0 (
raw/,wiki/,index.md,log.mdper the Karpathy method). Memory compounds from the first message even while automation is absent. - Safety keyword escalation live from day one.
Exit criteria (measured over pilot weeks 5–8)
| Criterion | Bar |
|---|---|
| Clients onboarded and active (≥ 4 interactions/week) | 10 / 10 onboarded, ≥ 8 active |
| Meal-photo behavior exists without automation | ≥ 60% of days with ≥ 2 photos |
| Draft quality baseline established | Edit-rate dashboard live; ≥ 60% of drafts approved unedited |
| Approval queue is fast enough to be sustainable | Median approval < 30 s/item; Nata total ≤ 60 min/client/week (down from ~110) |
| Coaching stays in-channel | ≥ 90% of client interactions via @NataCoachBot, not Nata's personal DMs |
| Safety rail fires correctly | 100% of seeded + real pain/injury phrases escalated |
2. Phase 1 — "Automation core" (Nov 2026–Jan 2027)
Thesis being tested: automation raises adherence and quality while cutting Nata's time ~4×.
Scope
- Health Sync live: Whoop API v2 + Garmin Health API, OAuth deep link from chat, webhook ingestion. Screenshots retired.
- Meal Lens live end-to-end: photo → macros → coach-grade feedback vs. Nata's food program, propose-confirm on macros, < 60 s photo-to-feedback.
- Session Mode v1 in chat: guided exercise-by-exercise runner with buttons; weight-proposal engine (history + today's recovery, progressive overload); RPE capture per exercise.
- Morning Brief live: recovery-driven daily plan proposal before 07:30 local.
- Graduated autonomy begins: message types move to
auto-send with auditvia the per-client streak mechanics in 09 §3.3 (20 consecutive unedited approvals per client, promotion confirmed by Nata); fleet-level we expect ≥ 90% approve-unedited over a rolling 100 items before most clients hit their streaks (expected first: meal feedback on green-recovery days, session confirmations, routine nudges). Program changes and anything safety-adjacent stay atdraft/escalate. - Coach Console v1: roster health view, alert feed, autonomy dial controls, cost telemetry.
Exit criteria — this phase owns the pilot success metrics:
| Criterion | Bar |
|---|---|
| Training adherence | ≥ 80% of planned sessions |
| Meal-log coverage | ≥ 70% of days (≥ 2 meals via Meal Lens) |
| Nata time per client | ≤ 30 min/week, instrumented |
| Retention | ≥ 8 / 10 opt to continue past week 8 of this phase |
| Proposal acceptance (weights, macros, adjustments) | ≥ 70% confirmed unedited |
≥ 2 message types running at auto-send with audit |
With < 5% post-hoc correction rate |
| Health Sync freshness | Recovery data present for ≥ 90% of Morning Briefs by 07:30 |
| Unit cost | LLM spend < $2/user/day |
3. Phase 2 — "Rich experiences" (Feb–Apr 2027)
Thesis being tested: rich surfaces deepen the relationship (not just decorate it), and Form Check works without custom CV models.
Scope
- Mini App workout runner: Next.js Telegram Mini App — live set/rest timers, plate math, one-thumb logging; chat runner remains as fallback (chat-first principle).
- Form Check live: video upload → vision-LLM technique analysis against per-exercise checklists → 2–3 coaching cues, Nata-reviewed at
draftuntil it earns its own dial movement; hard cases escalate with timestamped clips. - Weekly Review: Sunday chart-rich summary for the client (Mini App) + cross-client digest for Nata.
- Coach Console v2: full analytics, program builder (replacing the sheet), chat takeover.
- Apple Health via bridge (e.g. Health Auto Export) for non-Whoop/Garmin clients.
- First-tier autonomy: message types move to
fully autonomousvia the per-client mechanics in 09 §3.3 (30 audited sends, 0 retractions, ≥ 14 days in audit); the fleet-level expectation is < 3% correction over 300 audited items before promotions become routine. Intake grows 10 → 25 clients (waitlist) to stress the queue before Phase 3.
Exit criteria
| Criterion | Bar |
|---|---|
| Mini App runner adoption | ≥ 60% of sessions run in Mini App (rest in chat) — with session completion no worse than chat-only |
| Form Check usefulness | Nata agrees with the top cue on ≥ 75% of reviewed videos; ≥ 1 video/client/month uploaded |
| Weekly Review engagement | ≥ 70% of clients open it within 24 h |
| Autonomy | ≥ 2 message types fully autonomous with < 3% correction rate |
| Scale rehearsal | 25 active clients; Nata time ≤ 20 min/client/week; adherence and coverage bars from Phase 1 still met |
| Perceived personalness | ≥ 8/10 clients answer "feels like Nata is coaching me" ≥ 4/5 in monthly survey |
4. Phase 3 — "Scale to 100+" (May–Aug 2027)
Thesis being tested: the economics — one coach, 100+ clients, people paying, relationship intact.
Scope
- Autonomy dial mostly open: everyday coaching (meal feedback, session coaching, briefs, reviews)
fully autonomouswith sampled audits (Nata reviews a random 5%);draftreserved for program changes and plateau calls; safety alwaysescalate. - Squads: optional small group chats (5–8 clients, e.g. "postpartum strength", "fat-loss travelers") for community and challenges — the optional later from locked decision 1; 1:1 private chat remains the coaching channel.
- Payments: subscription billing at the target $79–129/month (provider chosen then — Stripe or Telegram-native; deliberately out of scope until now per brief §6).
- Ops hardening: rate-limit-aware queues, per-user cost caps, on-call alerting, data-export/delete self-service, Console triage views built for a 100-roster (exception-based, not per-client).
- Growth 25 → 100+ via waitlist and squad referrals.
Exit criteria
| Criterion | Bar |
|---|---|
| Roster | ≥ 100 active paying clients |
| Nata time | ≤ 10 min/client/week at the 100-client mark |
| Retention | ≥ 85% month-over-month across the paid base |
| Adherence / coverage | Phase 1 bars (80% / 70%) hold at 100+ clients |
| Economics | Gross margin per client positive with LLM < $2/user/day; payment conversion from pilot cohort ≥ 70% |
| Trust under autonomy | Complaint/correction rate on autonomous messages < 2%; safety SLA still 100% |
5. Timeline
gantt title NataCoach roadmap (planning assumption, start 2026-09-01) dateFormat YYYY-MM-DD axisFormat %b %Y section Phase 0 Concierge pilot Bot onboarding and approval queue :p0a, 2026-09-01, 21d Seed Wiki Brains and programs :p0b, 2026-09-08, 14d Run 10-client concierge pilot :p0c, 2026-09-22, 34d Gate review and edit-rate baseline :p0d, 2026-10-26, 6d section Phase 1 Automation core Health Sync Whoop and Garmin :p1a, 2026-11-01, 35d Meal Lens end to end :p1b, 2026-11-01, 42d Session Mode chat runner :p1c, 2026-11-17, 42d Morning Brief and proposal engine :p1d, 2026-12-08, 28d Graduated autonomy rollout :p1e, 2027-01-05, 26d section Phase 2 Rich experiences Mini App workout runner :p2a, 2027-02-01, 42d Form Check pipeline :p2b, 2027-02-15, 42d Coach Console full analytics :p2c, 2027-03-01, 45d Weekly Review and 25-client intake :p2d, 2027-03-16, 45d section Phase 3 Scale to 100 plus Autonomy mostly open :p3a, 2027-05-01, 45d Squads and community :p3b, 2027-06-01, 45d Payments and billing :p3c, 2027-06-15, 45d Grow roster to 100 plus :p3d, 2027-05-01, 122d
6. Riskiest assumptions, and how each phase tests them
Ordered by how fatal being wrong would be. "Kill/pivot signal" is the number at which we stop pushing and change the design instead.
| # | Assumption | If it is false | How the phases test it | Kill / pivot signal |
|---|---|---|---|---|
| 1 | Clients will photograph meals and tap ✓ consistently when friction is near zero | Meal Lens and the food program are decorative; half the product's data loop dies | P0 measures raw willingness with a hand-driven pipeline (bar: 60% of days); P1 must lift it to 70% with instant feedback; P3 must hold it at 100 clients | Coverage < 50% in P0 despite < 15 s per meal → pivot to voice/text meal capture or drop per-meal scoring for daily summaries |
| 2 | An LLM with the Wiki Brain + Nata's programs can draft feedback Nata sends mostly unedited | Coach-in-the-loop never graduates; Nata stays the bottleneck; economics collapse | P0 establishes the edit-rate baseline (bar: 60% unedited); P1 requires 90% for dial movement; P2/P3 require sustained < 3% / < 2% correction under autonomy | Unedited rate stuck < 75% after P1 prompt/persona iteration → keep humans in loop permanently and re-price, or narrow autonomous scope to confirmations only |
| 3 | Nata's time can compress ~110 → 30 → 10 min/client/week without clients feeling the difference | The product is a nicer workflow tool, not a scaling engine | P0 instruments her time (bar: ≤ 60); P1 gates on ≤ 30; P2 adds the perceived-personalness survey (≥ 4/5); P3 gates on ≤ 10 at 100 clients | Time plateaus > 40 min/client or personalness < 3.5/5 → cap roster ~30 and reposition as premium hybrid coaching |
| 4 | Recovery-aware proactive coaching improves adherence (vs. annoying people) | Morning Brief is spam; wearable integration is a gimmick | P0 ships no integration (baseline adherence + screenshot pain); P1 compares adherence and Brief response rates before/after Health Sync; mute-rate tracked | Brief response < 40% or mute rate > 20% in P1 → reduce to 2–3 proactive moments/week, triggered only on red-flag recovery |
| 5 | Vision LLM + per-exercise checklists give useful form feedback without custom pose models | Form Check ships embarrassing cues or nothing | P2 gates on Nata agreeing with the top cue ≥ 75%; disagreements become checklist edits, logged weekly | Agreement < 60% after two checklist iterations → Form Check becomes "video routed to Nata with AI pre-notes", revisit pose models post-P3 |
| 6 | Unit economics work: LLM < $2/user/day and clients pay $79–129/month | Scaling multiplies losses | P1 instruments per-user cost with model routing (Opus-tier reasoning, Haiku-tier routing); P3 tests willingness to pay (conversion ≥ 70% from pilot cohort) | Cost > $3/user/day after routing/caching work, or conversion < 50% → cut always-on surfaces, raise price, or gate expensive pipelines to higher tiers |
| 7 | Trust survives autonomy — clients accept that everyday messages are AI, backed by Nata | Dial-opening triggers churn exactly when economics need it | P1 audits every auto-sent message; P2 survey after first autonomous types; P3 watches complaint rate (< 2%) and retention through full opening |
Retention dips > 10 pts within a month of a dial change → roll the dial back one notch and re-earn with a longer audit window |
| 8 | Whoop/Garmin APIs are reliable enough for a 07:30 Morning Brief | The flagship proactive moment misfires on stale data | P1 gates on 90% data freshness; degraded mode ("no sync yet — how did you sleep?") specified in Health Sync | Freshness < 80% sustained → decouple Brief from sync (send on schedule, enrich when data lands) |
The vision this plan serves is in 01 · Product Vision; what each shipped phase feels like for Marta and Denys is in 02 · User Experience and 10 · Personas.