1. The problem: great coaching does not scale
A great personal coach is not an information product. The information — sets, reps, macros — has been free on the internet for twenty years. What clients actually pay for is a relationship: someone who knows their history, notices when they slip, adjusts the plan before Monday goes wrong, and whom they do not want to disappoint.
That relationship is expensive to produce. Nata's 1:1 remote coaching today (numbers are her real workload, lightly rounded):
| Activity per client | Time per week |
|---|---|
| Reading check-in messages, replying | ~45 min |
| Reviewing food photos sent over DM, replying meal by meal | ~30 min |
| Adjusting the week's program (sleep, travel, soreness) | ~15 min |
| Weekly summary + plan for next week | ~20 min |
| Total | ~110 min/client/week |
At ~110 minutes per client per week, a full-time coach saturates at 12–15 clients. Nata is at 14 and turning people away. Raising prices filters people out; hiring junior coaches dilutes exactly the thing clients came for — her.
The industry's answer — generic fitness apps — solves the wrong problem. Apps scale content, not relationship:
- Nobody is watching. The app does not care if you skip Tuesday. Typical consumer fitness apps lose the majority of users within the first month (industry-typical 30-day retention is ~25–30%; assumption, directionally well-established).
- The user does the work. Logging food, choosing weights, deciding what to do on a bad-sleep day — the app pushes labor onto the person least equipped to do it consistently.
- No memory that compounds. The app on day 300 knows barely more about you than on day 3.
So the market splits into two bad options: a $15/month app that ignores you, or a $300/month human who cannot take you. NataCoach is built for the gap between them.
2. Who Nata is, and what only she brings
Nata is a real strength and nutrition coach (9 years coaching, ~200 clients over her career; current specialties match our two pilot personas: postpartum return-to-strength and fat loss for desk-bound professionals — see Marta & Denys). Her clients stay for years, and when asked why, they do not say "the programs." They say: "She notices."
What only Nata brings, and what the system must never fake:
- Programming judgment. Turning "8 months postpartum, wants a pull-up, sleeps badly" into a periodized 12-week program is expert work. The AI executes and adapts programs; it does not author them.
- Judgment at the edges. Pain that might be an injury, a plateau that might be underfeeding, a client quietly spiraling — pattern recognition from a decade of humans, not a dataset.
- Presence. A real voice note from Nata after your first full pull-up negative is worth fifty perfectly worded bot messages. Her authentic presence is the moat, and we spend it deliberately (see Coach-in-the-loop).
- Permission to be trusted. People accept "cut today's volume 20%" from someone they believe has earned the right to say it.
3. The product thesis
AI handles the every-day surface area. Nata supplies programs, judgment, and presence.
Roughly 80% of Nata's 110 weekly minutes per client is surface area: reading, acknowledging, computing, proposing, reminding, summarizing. Almost none of it requires her judgment — but all of it requires someone, every day, or the relationship decays. That surface area is exactly what an LLM with a persistent per-client memory (the Wiki Brain) does well.
| Every-day surface area → @NataCoachBot | Judgment & presence → Nata |
|---|---|
| Morning Brief from last night's recovery data | Authoring training + food programs |
| Proposing weights for every set (Session Mode) | Approving/editing AI drafts while autonomy is low |
| Meal photo → macros → feedback (Meal Lens) | Program changes, plateau calls, injury triage |
| Logging everything as events; nightly Wiki Brain distillation | Occasional real voice notes at moments that matter |
| Nudges, confirmations, Weekly Review draft | Reviewing the cross-client Coach Console digest |
Target end state: Nata's time drops from ~110 to ≤ 30 min/client/week in the pilot (and ≤ 10 at scale), while the client experiences more daily contact than a 1:1 human coach could ever give.
4. The seven locked decisions — deliberate improvements over the original ask
The original request (brief §2) was directionally right. Where we changed it, we changed it on purpose:
| # | Original ask / idea | Locked decision | Why it is an improvement |
|---|---|---|---|
| 1 | "Subfolders / different chats with different people, adding the bot to each" | One bot, N private chats. Every client simply DMs @NataCoachBot; a Telegram bot automatically has a separate private chat with each user who starts it. No folders, no per-person setup, no group chats for coaching. | Folders are a feature of human Telegram accounts, not bots — the proposed mechanic does not exist to build. Adding a bot to per-person group chats would leak a second member into every "private" coaching thread, multiply setup per client, and cap scale. Nata's per-client view belongs in the Coach Console, where it can show charts and queues Telegram never could. Squad community chats come later as an optional extra, not the coaching channel. |
| 2 | Implied: a bot, maybe a separate app for rich features | Chat-first, Mini App for rich moments. Everything works in plain chat; a Telegram Mini App opens for the workout runner, weekly charts, and the Coach Console. | Zero installs, zero new logins. The client never leaves the app they already open 20× a day — which is the whole distribution thesis of building on Telegram. |
| 3 | User starts a training, uploads photos, asks questions | Proactive, recovery-aware coaching. The bot opens conversations: Morning Brief from Whoop/Garmin recovery, pre-workout nudge, instant meal feedback, Weekly Review. | A reactive bot is a generic app with extra steps. Accountability means the system notices first. "Rough sleep — cutting today's volume 20%, ok?" before breakfast is the product. |
| 4 | "We propose the weights… we propose macros" (stated as an example) | Propose-confirm everywhere — elevated from example to the universal interaction contract: weights, macros, weigh-ins, schedule changes. One tap confirms; editing is the exception. | Every logging burden the user carries is a churn vector. If the default action is ✓, adherence stops depending on willpower. |
| 5 | Not in the original ask | Coach-in-the-loop with graduated autonomy. LLM drafts; Nata approves from a queue (< 30 s/item); per-message-type dial draft → auto-send with audit → fully autonomous. Pain/injury/medical always escalates. |
This is the mechanism that makes "AI coaching in Nata's name" honest. Day one, every word is hers by approval; autonomy is earned per message type with measured edit rates, never assumed. |
| 6 | "Personal LLM wiki brain (Karpathy method)" | Event-sourced truth, distilled memory. Immutable events in Postgres (+ raw/ in the wiki repo); nightly distillation maintains wiki/, index.md, log.md. Analytics are deterministic SQL over events, never LLM output. |
Faithful to the Karpathy method (raw sources the LLM reads but never edits) and it makes Nata's dashboards trustworthy: a chart of adherence is arithmetic, not a hallucination. See System Architecture. |
| 7 | Not in the original ask | Safety rails. Pain/injury keywords pause programming and escalate; macros labeled as estimates; contraindications captured at onboarding; medical questions deflected to professionals; strict health-data privacy. | A coaching product that touches postpartum training and health data has non-negotiable failure modes. Rails are designed in from the first message, not patched in after the first incident. |
5. Product principles
- The user does less. The prime directive. Every flow is scored by taps required and seconds spent. If a feature adds user labor, it is wrong until proven otherwise. Target: an engaged day costs the client under 3 minutes of interaction.
- Chat-first. Plain Telegram chat — photos, buttons, voice — is the product. The Mini App is for moments a chat genuinely cannot serve (live workout runner, charts, Console). If it can be a message with two buttons, it is a message with two buttons.
- Propose-confirm. The system computes the default; the human taps ✓ or adjusts. This applies to clients (weights, macros, weigh-ins) and to Nata (drafted feedback, suggested program tweaks). Nobody starts from a blank input field.
- Coach-in-the-loop. Autonomy is a dial, not a switch, set per message type and moved only when edit-rate data says the AI has earned it. Safety-critical messages never leave
escalate. - Compounding memory. Every interaction makes the next one better, because everything distills into the Wiki Brain. Month 6 must feel noticeably more personal than week 1 — that is the retention engine and the defensibility.
- Safety first. When in doubt: pause the plan, escalate to Nata, say less. Estimates are labeled as estimates. The bot never diagnoses.
6. The core value loop
The loop everything else in these docs serves:
flowchart LR A["Data in<br/>one-tap confirms, meal photos,<br/>wearable sync, chat"] --> B["Wiki Brain<br/>distilled, current model<br/>of this client"] B --> C["Better coaching<br/>recovery-aware, specific,<br/>in Nata's voice"] C --> D["More trust<br/>advice lands, effort feels seen,<br/>confirming stays effortless"] D --> A N["Nata<br/>programs, edits, voice notes"] --> B N --> C
Generic apps never close this loop: data goes in, nothing personal comes back, logging stops by week 3. NataCoach closes it daily — and Nata's edits inject ground truth the loop could never learn on its own.
What the loop feels like on a random Tuesday (Marta, week 5 — full transcripts in User Experience):
07:10 @NataCoachBot: Morning, Marta ☀️ Whoop says recovery 44% — rough
night (2 wake-ups, 5h 40m sleep). Today was Pull B. I'd keep the
pull-up negatives but drop rows from 4 sets to 3 and skip the
finisher. Sound right?
[ ✓ Do that ] [ Keep full plan ] [ Move to tomorrow ]
07:11 Marta: ✓ Do that
07:11 @NataCoachBot: Done — 35 min instead of 50. See you at the bar 💪
Two taps. The plan adapted, the event was logged, the Wiki Brain got a little sharper about how Marta's bad nights look, and Nata spent zero minutes.
7. Positioning
| Dimension | Generic fitness app | 1:1 human coach (Nata today) | NataCoach |
|---|---|---|---|
| Price (monthly, assumption) | $10–30 | $250–350 | $79–129 target |
| Personalization | Template picked from a quiz; static | Fully personal, limited by coach hours | Nata-authored program + Wiki Brain adapting it daily |
| Accountability | Push notifications into the void | Real, but gated on when the coach checks messages | Proactive daily contact within minutes, coach-backed |
| Reaction time to your day | None — the plan does not know you slept badly | Hours, if you think to mention it | Morning Brief before breakfast, automatically |
| Memory | Session logs; no synthesis | In the coach's head; lossy at 14 clients | Wiki Brain: compounding, inspectable, Nata-editable |
| Scale | Millions of users, zero relationship | 12–15 clients, deep relationship | 10 → 100+ clients per coach, relationship preserved |
| Safety handling | Generic disclaimers | Good, when the coach sees the message | Keyword escalation + contraindication rules, always-on |
The wedge: coach-grade personal attention at roughly one-third the price of 1:1 coaching, at 10× a coach's capacity — a segment neither incumbent can serve.
8. Pilot success metrics (10 users, 8 weeks)
All metrics are computed deterministically from events (locked decision 6), visible live in the Coach Console. Targets are honest bars, not vanity:
| Metric | Definition | Target | Why this bar |
|---|---|---|---|
| Training adherence | Planned sessions completed (incl. bot-adjusted versions) / planned sessions | ≥ 80% | Nata's 1:1 clients average ~85% (assumption); a scalable system may cost a little, not a lot |
| Meal-log coverage | Days with ≥ 2 meals through Meal Lens / total days | ≥ 70% of days | Self-logging apps see coverage collapse < 30% by week 4 (assumption); photo + propose-confirm must beat that decisively |
| Nata time per client | Console + approval-queue + chat-takeover time, instrumented | ≤ 30 min/client/week | ~4× compression vs. her ~110 min today; the economic thesis in one number |
| Retention | Clients active in week 8 who opt to continue past the pilot | ≥ 8 / 10 | Below this, the relationship is not surviving the automation |
| Proposal acceptance | Proposals confirmed unedited (weights, macros, adjustments) | ≥ 70% | Measures whether propose-confirm defaults are actually good |
| Draft approval rate | Nata approves AI drafts without edit | ≥ 80% by week 8 | The gate for moving the autonomy dial in the roadmap |
| Safety escalation SLA | Pain/injury flags acknowledged by Nata | 100% < 60 min (waking hours) | Non-negotiable |
| Unit cost | LLM spend per user per day | < $2 | Locked cost target (brief §6) |
Secondary signals we watch but do not gate on: median user taps/day (expect 6–10), Morning Brief response rate, voice-note reactions, unprompted messages to the bot (a trust proxy — we want these to grow).
9. Non-goals (for now)
Per the locked technical decisions (brief §6): no custom pose-estimation models (vision LLM + checklists first — see Form Check), no barcode scanning, no native iOS/Android apps, no payments until Phase 3, no multi-coach marketplace. NataCoach scales Nata; generalizing to "any coach" is a question for after the pilot proves the loop.
Next: how this vision becomes daily experience in 02 · User Experience, and how it gets built in stages in 11 · Roadmap.