← All capstone projects

Sports

PitchingIQ

Built by Sachkouskaya Cohort 9 Sports performance / baseball training analytics

PitchingIQ is an AI performance product for high school and college pitchers that connects wearable biometrics with pitching-specific training and game outcomes. Instead of treating a generic recovery score as universally meaningful, it looks for player-specific patterns across sleep, strain, velocity, command, fatigue, and pitch count. The product delivers verified patterns and monthly reports meant to help athletes train smarter and improve recruiting readiness.

The problem

High school and college pitchers train hard but fly blind — no idea why some days they're sharp and others flat. Elite systems like TrackMan and motion capture cost $15,000–$25,000 and only reveal what's mechanically wrong, never why it's happening from a pitcher's own habits and physiology. The workload-to-velocity link is established in research, but only at the population level, ignoring recovery and sleep and missing non-game throwing volume. So no tool answers the individual question: does this pitcher's velocity drop when he sleeps poorly three nights before a start, or when recovery is low? The same biometric reading means different things for different pitchers.

The solution

PitchIQ is a personal pattern-discovery engine — not a coaching or training platform — that connects a pitcher's own wearable biometrics to their training and game performance and surfaces recurring patterns confirmed from that athlete's history alone, never population averages. The differentiators are individual-only baselines, biometric-plus-performance correlation that competitors don't attempt, and repetition-confirmed patterns that surface only after a pattern recurs enough times in that athlete's record — the core method IP. It delivers $25K-system-level insight from data the athlete already collects, at $15/month, working three ways: on-demand pattern checks, an agent that continuously finds and confirms patterns, and a monthly self-comparison report.

How it works

The entire experience runs inside a Telegram chatbot. The athlete logs each day via structured prompts (rest/training/game), while WHOOP biometrics — HRV, recovery, resting heart rate, sleep efficiency, strain — fetch automatically via API. A Pattern Agent tests preselected metric pairs on a cadence using Spearman correlation plus a repetition-confirmation gate; only patterns passing that gate are pushed to the athlete for ✓/✗ verification, and verified patterns are stored permanently. Three models each handle a task: Claude for pattern interpretation, on-demand answers, and report generation; Whisper for voice-note transcription; and OpenAI text-embedding-3-small for a RAG science layer that cross-checks each pattern against six peer-reviewed articles as supported, contradicted, or no evidence. A strict hallucination guardrail ensures the LLM narrates only stats-validated patterns and never invents correlations, hedging every insight to the data available. On 180 days of synthetic data the agent surfaced 3 of 4 embedded patterns, confirming a ~50-day-per-day-type minimum.

Who it's for

The primary customer is B2C: the individual pitcher. The persona is "Marcus," a 17-year-old high school right-hander throwing ~88 mph, chasing D1 attention before signing day, who already owns a WHOOP, logs willingly, and lives on his phone — with parents paying the $15/month subscription. College pitchers form a secondary segment motivated by roster security. Phase 2 adds B2B2C institutional tiers for academies, travel organizations, and small college programs (~$4K/year). A firm trust principle governs all phases: institutions never see raw biometrics — coaches see only confirmed patterns and the readiness signal — because a pitcher who fears being benched over his data disengages, destroying the dataset that is the moat.

Why it matters

AI-powered athlete performance analytics sits at roughly a 20–25% CAGR, within a large and growing US high school athlete population (8.26M+ in 2024–25). Comparable biometric platforms are richly funded — WHOOP raised $575M at a $10.1B valuation, Oura reached ~$11B — yet all aim at pros, elite teams, or broad consumer health. None serve the individual amateur pitcher. Bottom-up, the US pitcher beachhead is a ~$28M TAM, extending toward $300–500M across adjacent individual sports. A pre-seed startup with a working end-to-end prototype, PitchIQ's next step is a closed pilot with one real pitcher over a full season to prove the method on live data. The compounding, per-athlete correlation dataset — never sold — is the long-term defensibility that late entrants can't replicate.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Olga Sachkosukaya
Your Product:PitchIQ
Your Industry:AI-powered athlete performance intelligence
Date:June 6, 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?PitchIQ helps high school and college pitchers train smarter and get recruited by turning their wearable, training and game performance data into their own personal performance patterns. Industry: AI-powered athlete performance intelligence — specifically the individual-athlete layer of the AI-in-sports market, at the intersection of wearable biometrics and sport-specific performance analytics. What it is: a personal pattern-discovery engine — not a coaching or training platform. For one pitcher, it connects that athlete's own wearable biometrics (HRV, sleep, strain, recovery %) to their own training and game performance, and surfaces recurring patterns confirmed from their history alone — never population averages. Joe at 12% recovery threw great at his 4pm game; James at 12% recovery struggled every time — pitchIQ treats these as different signals, because they are. It works three ways: on-demand pattern checks the athlete requests, an agent that continuously finds and confirms recurring patterns, and a monthly self-comparison report. Problem we solve: Elite systems (TrackMan, motion capture) cost $15,000–$25,000 and tell a pitcher what is mechanically wrong — never why it's happening from his own habits and physiology. The workload-to-velocity link is already established in scientific research (Romano & Etim-Andy, 2023), but only at the population level, ignoring recovery and sleep and missing non-game throwing volume entirely — so it tells an individual pitcher little about himself. No tool answers, for this pitcher: does my velocity drop when I sleep poorly three nights before a start? when recovery is low? pitchIQ surfaces those personal, confirmed patterns from data the athlete already collects, at near-zero marginal cost. Who it's for: initial target audience is high school and small-college pitchers (and the parents who fund them) who lack elite analysis infrastructure but already own a wearable and want a personal edge toward recruitment.Please leave this area blank. This space is for the Instructor to provide you with feedback.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?WHY NOW (growth drivers) Sports analytics is growing fast — ~15–25% CAGR depending on source (MarketsandMarkets: $2.29B in 2025 → $4.75B by 2030, 15.7%; Fortune: 20.5% through 2034). Wearables are shifting from tracking to prediction — demand for subscription-based health/performance intelligence is proven and accelerating (see market signals below). Teams want injury prevention and training optimization — the established enterprise use case (Catapult). The athlete population is large and growing — U.S. high school sports participation hit a record 8.26M+ athletes in 2024–25 (NFHS). The under-built corner: AI capital and tooling are abundant (AI took 34%+ of global VC in 2025), but almost none points at the individual sub-elite athlete — the segment every major player skips. MARKET SIGNALS (comparable companies — category is fundable and growing) WHOOP — raised $575M Series G at a $10.1B valuation (≈3× its prior $3.6B); 2.5M+ members, 103% YoY bookings growth, $1.1B run rate, cash-flow positive in 2025. Source: businesswire / Bloomberg Oura — reached an ~$11B valuation in late 2025 — a second subscription-biometric platform at scale. Source: Bloomberg Catapult Sports — FY2025 revenue $116.5M, up 16.5%; wearable tracking and analytics for elite teams. Source: ASX:CAT filings Hudl — sports performance/video analytics; Bain Capital invested to accelerate global growth. Kitman Labs — "performance intelligence" platform for elite sport and defense. Read: huge appetite for performance/biometric intelligence — all aimed at pros, elite teams, or broad consumer health. None serve the individual amateur pitcher. MARKET SIZE Top-down context (not pitchIQ's addressable market): Global sports technology: $34.25B in 2025 → $68.71B by 2030 (14.9% CAGR). Source: MarketsandMarkets Global sports analytics (core): $5.79B in 2025 → $31.14B by 2034 (20.5% CAGR). Source: Fortune Business Insights Performance-focused analytics (wearables, training optimization, injury prevention): ~$1.5B–$3.5B today. Bottom-up (pitchIQ's actual plan — price $15/month, ≈$180/yr per athlete): Layer 1 — U.S. baseball pitchers (beachhead): ~155,000 pitchers × $180 ≈ $28M TAM Layer 2 — add U.S. baseball position players: ~365,000 more → baseball-wide ≈ $90M TAM Layer 3 — adjacent individual sports (tennis, golf, track…): platform extension → $300–500M TAM (directional) SAM (serviceable now — wearable-owning, reachable U.S. pitchers, 15–20%): ~23,000–31,000 × $180 ≈ $4–5.5M SOM (realistic 5-year): path to ~$1.5–2M ARR Convergence check: the bottom-up multi-sport layer (~$300–500M) lands in the same range as the top-down U.S. amateur beachhead estimate (~$150–500M). HEADWINDS Cold-start: confirmed patterns need several weeks/months of the athlete's own data. Wearable-API dependency (Whoop): platform risk if terms change. Mitigated by device-agnostic architecture (extensible to Oura, etc.). Privacy/consent for under-18 biometric data — a real compliance requirement built in from the start. Academy sales cycles are long — requires a structured validation-partnership approach. Method must be proven on real pitcher data before institutional scale — validation precedes scale by design. High schoolers have limited direct purchasing power — parent-funded most of the time. KEY COMPETITORS — AND WHY THEY DON'T SOLVE THIS Motus Global — arm-sleeve workload; mechanical load only, no biometrics. Driveline PULSE — throw-load tracker; no recovery layer, no personal correlation. Catapult — enterprise team monitoring; population benchmarks, priced for pro/Power 5. TrackMan — $15K–$25K radar; mechanics only, zero physiology. Whoop Coach (OpenAI-powered) — general fitness coaching over Whoop data; not sport-specific, doesn't ingest pitching performance or confirm correlations between performance and biometrics. Uplift Labs — motion capture for pro/NCAA; biomechanics only, no biometric layer.
What is the projected growth rate of your target market segment over the next 3-5 years?Projected segment CAGR: ~20–25% (2025–2030). pitchIQ's segment — AI-powered athlete performance analytics — isn't sized on its own by analysts, so the range is triangulated from the adjacent markets that bound it: Market segment Projected CAGR Role Broad sports technology ~14.9% Floor / context Sports wearables ~15% Addressable-base driver AI-in-sports ~21–28% Closest proxy Sports-tech analytics >29% Ceiling (directional — skewed to team/fan analytics) pitchIQ target segment ~20–25% Triangulated estimate The real driver for pitchIQ is adoption, not market CAGR: the athlete population is large but roughly flat (~8.26M U.S. high school athletes), so the base expands chiefly as wearable ownership among young athletes rises — segment growth ≈ category CAGR (~20%+) compounded by accelerating wearable penetration. Our own obtainable share of that growing segment (SOM): Year Individual subs Orgs / Academies ARR Year 1 100 (word of mouth, HS pitchers) ~$18K Year 2 500 1 academy (first sign) ~$95K Year 3 2,000 ~10 orgs + 3 academies ~$410K Year 5 7,000 ~80 orgs + 15 academies ~$1.6M
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Stage 0: early-stage startup — prototype / pre-seed phase (pre-revenue, pre-validation). We have a working end-to-end prototype: a Telegram chatbot collects training and game performance logs through a structured input interface, automatically pulls WHOOP biometric data via API, and stores both on Google Drive. The bot issues a monthly report, answers on-demand questions about the athlete's patterns, and runs automated pattern searches. Next stages: 1. Validated MVP (next ~3–6 months) — accumulate 2–3 months of real data from the first athlete, then refine and prove the pattern method on real inputs, not just synthetic. 2. Early traction / seed (Year 1–2) — expand to 2–5 athletes to stress-test the method; convert first paying individual users (HS pitchers, word of mouth); sign first academy in Year 2. 3. Scale-up (Year 3+) — grow individual subscribers and academy/org contracts; begin platform expansion (hitters, then adjacent individual sports). 4. Mature — established multi-sport platform with a compounding, defensible proprietary dataset.
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)What we sell: a pitcher's own performance intelligence — a personal engine that connects their wearable biometrics to their training and game data and surfaces confirmed individual patterns. We sell AI powered insight, not hardware, coaching, or data. Primary model: recurring subscription (SaaS). B2C individual (Phase 1, core): HS and small-college pitchers at $15/mo ($180/yr); word-of-mouth growth. B2B2C institutional (Phase 2, Year 2+): academies, travel orgs, small colleges license a team tier (~$4K/yr, TBD); coaches see only the readiness signal, never raw biometrics. Freemium layer (to test): free basic logging + limited summary to drive adoption; converts to paid once enough history unlocks confirmed patterns — the cold-start period doubles as the free trial. Why subscription: value compounds with continuous data (patterns sharpen over months), and recurring spend matches how travel-baseball families already pay. Not: transactional, marketplace, data resale, or hardware margin. The proprietary correlation dataset deepens the moat but is never sold — it's what makes pitchIQ an acquisition target. (Firm: $15/mo individual. Assumptions to lock: institutional price, freemium vs. paid trial.)
Who is your primary customer base (B2B, B2C, B2B2C)?Primary — B2C: the individual pitcher. HS pitchers (15–18): showcase/recruiting-driven, Whoop owners, scholarship stakes create urgency. College pitchers (18–22, NCAA D1–D3, NJCAA): motivated by performance and roster security. Own and control their own Whoop data; pitchIQ connects through the player's personal account. Player (or parent) pays the monthly/annual subscription. Phase 2 — B2B2C (institutional, Year 2+), in priority order: 2A Private academies (validation anchor): IMG Academy, Driveline, Cressey, premium regional academies. First because performance directors already understand biometric data — shortest validation path; one IMG partnership beats any marketing spend. Institutional budget, annual contract, highest data quality. 2B Travel organizations: Canes, East Cobb, USA Prime, Perfect Game programs. 10–20 pitchers per team across multiple teams; high-investment families. Natural land-and-expand — the same pitcher can be a Phase 1 subscriber whose org later buys a team tier. 2C Small college programs (NJCAA, D3, community colleges): slowest cycle, lowest budget, but strongest mission narrative for investors. Trust principle (all phases): institutions never get raw biometric data. Coaches see only confirmed patterns / the readiness summary — never HRV, sleep, or recovery values. Players control what they share and can revoke anytime. This is a business necessity, not just ethics: a pitcher who fears being benched over their data disengages, destroying the personal dataset that is pitchIQ's core moat. Trust = retention = dataset growth = defensibility.
DifferentiatorsWhat are the key differentiators for your company?Individual-only baselines — compares a pitcher only to their own history, never population averages. Same biometric reading means different things for different pitchers (Joe vs. James). Biometric + performance correlation — connects wearable data (HRV, sleep, recovery) to actual pitching outcomes; competitors do mechanics or biometrics, never the link. Repetition-confirmed patterns — surfaces a pattern only after it recurs enough times in that athlete's record, filtering false signals (core method IP). Accessible to the individual amateur — delivers $25K-system-level insight to HS/small-college pitchers from data they already collect, at $15/mo. Compounding proprietary dataset — every athlete-month deepens an individual correlation library that late entrants can't replicate — the long-term moat.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?N/A
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?N/A
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?N/A
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)Primary persona — "Marcus," 17, high school RHP Situation: junior, ~88 mph, plays travel ball (e.g., Canes-tier), a few D2 offers, chasing D1 attention before signing day; a Perfect Game showcase and key tournaments coming up. Goal: protect and raise velocity, show up ready for the outings that get him recruited. Pain: trains hard but flies blind — no idea why some days he's sharp and others flat; can't afford TrackMan/MoCap, and generic wearable advice isn't about his pitching. Behavior: already owns a WHOOP, logs/tracks willingly, lives on his phone (Telegram-friendly), data-curious but not technical. Pays: parents — already spending thousands/year on gear, coaching, showcases; $15/mo is trivial for any recruiting edge. Why pitchIQ: tells him what his own body and habits do to his performance, and gives him a data-backed story for recruiters. Secondary persona — college pitcher, 18–22 (NCAA D1–D3, NJCAA): motivated by roster-spot security and durability over a long season; often pays for himself.
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?The entire experience today runs inside a Telegram chatbot — no separate app. The athlete's full journey: -Opens the pitchIQ Telegram chatbot from a home-screen shortcut on his phone. -The chatbot asks the day type — he taps Rest / Training / Game. -The chatbot walks him through structured prompts for that day (pitch count, velo, command, how he felt, etc.) — buttons and numbers, no free typing. Under a minute. -The chatbot confirms and saves the entry; WHOOP data (HRV, sleep, recovery, strain) pulls in automatically in the background — he does nothing for that. -He can message the chatbot a point question about his own numbers/patterns anytime and get an answer back in the chat. -Once a month, the chatbot sends him a summary report in the chat — what went up, what went down, any signals it picked up. -He repeats the daily log; data keeps accumulating in the background.
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?Pain Points — friction in the current Telegram chatbot journey (prototype): 1.Daily logging discipline — the athlete must remember to log every day with no enforced reminder yet; missed days leave gaps in the data. (Most frequent) 2.Thin data, low payoff — with only ~2 weeks of history, the bot can't surface meaningful confirmed patterns, so the athlete puts in daily effort before the product can give much back. (Most severe) 3.Unproven, interim insights — pattern-finding runs on preselected pairs + Spearman and is validated only on synthetic data, so the answers don't yet feel trustworthy or individualized. 4.Manual effort with a slow feedback loop — logging is a daily ask while the main payoff (the monthly report) is far off, a poor fit for a teen wanting quick wins. 5.Fragile single-user prototype — no real onboarding, error handling, or recovery if a log or WHOOP pull fails; structured fields can't capture context the form doesn't cover.
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.AI Opportunity — pain points addressable with LLM-powered AI (ranked by severity/frequency) 1.Thin data, low payoff — LLM delivers value before patterns confirm: interprets short history in plain language, answers questions from day one, explains why more data is needed. 2.Daily logging discipline — conversational bot makes logging feel like a chat + sends contextual nudges, reducing missed days. 3.Unproven/interim insights — LLM turns validated correlations into clear, personal, hedged language and powers on-demand /ask (within the safeguards above). 4.Slow feedback loop — LLM auto-generates the monthly report and answers in-chat anytime, replacing one far-off payoff with continuous insight. 5.Limited structured input — LLM can later parse free-form/voice notes into structure, capturing context fixed fields miss.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.Diverge — idea generation (quantity over quality) 1.Conversational logging agent — log the day by chatting, not filling forms. 2.On-demand /ask — athlete asks any question about his own data, gets a plain-language answer. 3.Auto-narrated pattern alerts — LLM explains each statistically confirmed pattern in plain English. 4.Smart contextual nudges — reminders tied to recent activity ("logged a bullpen — how's the arm?"). 5.Auto-generated monthly report — LLM writes the recap (up/down, confirmed signals). 6.RAG science layer — cross-check each pattern against published sports-science literature. 7.Hallucination guardrail — LLM only narrates stats-validated patterns; never invents correlations. 8.Confidence/hedging engine — phrases every insight to the data available ("still confirming"). 9.Voice/free-text note parser — extract structure from "felt flat, tight shoulder" notes. 10.Pre-outing readiness summary — "based on your patterns, here's what helps you show up sharp." 11.Recruiter-ready data story — LLM drafts a shareable summary of the athlete's trends. 12.Coach view generator — LLM produces a patterns-only summary, no raw biometrics. 13.Memory-grounded personalization — answers anchored to the athlete's own stored history. 14.Anomaly explainer — flags and explains unusual readings vs. the athlete's baseline.
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.Converge — rank by impact × feasibility; pick the focus Top 3: -On-demand /ask, memory-grounded + hallucination-guarded (#2 + #7 + #13) — highest impact (delivers value from day one, the core engagement fix), feasible now with the existing pipeline. -Auto-narrated, science-checked pattern alerts (#3 + #6 + #8) — high impact (this is the product's payoff), moderate feasibility (needs RAG + confidence logic). -Conversational logging + smart nudges (#1 + #4) — solves logging discipline, very feasible, but lower differentiation.
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?The athlete logs daily data via the Telegram Chat Bot — a structured sequential form, <2 min. WHOOP biometrics fetch automatically overnight via a background agent. Mental notes can be added any time via /mental (free-form text or voice). Every 3 days the Pattern Agent tests preselected metric pairs and surfaces statistically significant correlations; only patterns passing a repetition-confirmation gate are batched and pushed to the athlete each evening at 8pm for verification (✓/✗). Verified patterns are stored permanently and included in all future reports. The athlete can query their own data on demand via /ask, and receives weekly (text) and monthly (text + PDF) reports. All LLM output narrates only stats-validated patterns — never invented ones — grounded in the athlete's own history and curated science.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?The primary interface is Telegram — no custom app screens. Interaction design is documented as a conversation flow built on 5 commands: /log — daily data entry, sequential questions + review step before saving /mental — free-form text or voice note, any time /ask — on-demand correlation/trend queries, returns Spearman ρ + a science reference /report — weekly (text) or monthly (text + PDF) summaries /patterns — review and verify pending patterns (✓/✗) AI outputs appear inline at key moments: a confidence score on each pattern push, ρ and p-value on /ask answers, and a science citation (or "no evidence") per pattern. Every insight is hedged to the data available. Bot Flow diagram: https://www.figma.com/board/xO5SB4mTIRSPazsQSBQIkJ/Pitcher-Log-%E2%80%94-Bot-Flow-v1.0?node-id=0-1&t=0pGwMnL43jpnV7IA-1 Database Schema: https://www.figma.com/board/ax9ZpxIEhHSZeRPq05SrzE/Pitcher-Log-%E2%80%94-Database-Schema-ER-Diagram?node-id=0-1&t=HTVKwvybS1NsnQX3-1
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?The prototype demonstrates 3 core AI interactions: (1) data collection via a Telegram bot simulation where the athlete logs a training day and the bot confirms save; (2) pattern discovery where the Pattern Agent loads the dataset, tests preselected metric pairs, and surfaces statistically significant correlations with ρ and p-values; (3) scientific validation where each pattern is cross-referenced via RAG against curated sports-science sources, shown as supported / contradicted / no_evidence. Note: the demo runs on representative synthetic data (180 days, sample pitcher) to show the full loop — not yet proof on a live athlete. Essential for launch: /log and /mental commands, WHOOP API integration, Pattern Agent (3-day cadence), pattern-verification push, weekly/monthly report as Telegram text, and the hallucination guardrail (LLM narrates only stats-validated patterns). Deferred to later releases: /ask on-demand queries, monthly PDF report, voice input for /mental, multi-athlete support. Prototype (Telegram chatbot demo): https://claude.ai/public/artifacts/643882af-52d0-42d9-adb5-7b50263b07cf
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?Tone: precise and data-first, no motivational filler. The AI speaks like a sports scientist, not a wellness app. System instruction: "You are a performance analytics assistant for a competitive baseball pitcher. Your role is to process the athlete's own data and surface statistically grounded insights from his individual history only — never population averages. Never speculate or invent correlations: report only patterns produced by the statistical layer (Spearman + repetition-confirmation gate). When referencing a pattern, always cite the correlation coefficient (ρ), p-value, strength, and supporting research status. State confirmed patterns directly; for anything below the confirmation threshold or built on limited data, say so explicitly and hedge accordingly. If data is insufficient, say so plainly rather than guessing." Input format: structured, tagged data blocks — daily log fields, WHOOP metrics, and mental-note transcripts passed separately with clear labels. Output format (consistent across responses): Pattern descriptions follow the template: [Metric A] → [Metric B] · ρ = X · [strength] · [science status] Report summaries use delta indicators: ↑ up · ↓ down · → no change Confirmed findings stated directly with their statistics; unconfirmed/low-data findings clearly flagged as provisional.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?Good output is defined by the following benchmarks: statistical validity (all surfaced patterns meet thresholds |ρ| > 0.25 and p < 0.05), pattern accuracy (correlation found on synthetic data matches embedded ground truth within ρ ±0.05), RAG relevance (science reference returned is topically relevant to the metric pair), hallucination avoidance (LLM does not reference metrics or correlations absent from the dataset), response clarity (athlete understands the pattern and its implication without additional explanation), and latency (/ask response delivered within 10 seconds, /mental LLM processing within 60 seconds).
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?Standard case: 180 days of synthetic data with 4 embedded patterns, Pattern Agent expected to surface at least 3 of 4 with correct direction (positive or negative ρ). Edge case — insufficient data: fewer than 30 training days logged, Agent expected to return "insufficient data" and not surface spurious correlations. Negative case — no real correlation: two unrelated metrics such as resting heart rate vs. strikeouts, Agent expected to discard the pair (p > 0.05) and surface nothing to the athlete. Mental log edge case: voice note with background noise and incomplete sentence, LLM expected to extract available signal, flag low confidence, and not hallucinate missing context.
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?PitchIQ uses three AI models for different tasks. Claude API (Anthropic) handles pattern interpretation, /ask responses, and report generation — chosen for its ability to reason over structured data and produce precise, citation-backed outputs without hallucinating. Whisper API (OpenAI) handles speech-to-text conversion for /mental voice notes — chosen for accuracy on short informal speech and simple API integration. OpenAI text-embedding-3-small handles vector embeddings for RAG — converts research articles into searchable vectors stored in pgvector, enabling semantic search across the scientific knowledge base. Main limitation: all 3 models require API calls, meaning the system depends on external services and incurs per-token costs that scale with usage. Integration: Claude and Whisper are called via Python scripts triggered by bot commands, embeddings are pre-generated and stored in PostgreSQL with pgvector extension.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.Daily log (via /log command): day type (rest / training / game), sourced from athlete input via Telegram bot, required to determine which question set to load. Training day fields: session type, pitch count, arm fatigue (1–10), command rating (1–10), velocity, plus 5 additional structured fields — all required for pattern analysis. Game day fields: innings pitched, strikeouts, earned runs, velocity, nerves (1–10), plus 10 additional structured fields. Whoop biometrics (fetched automatically): HRV, recovery score, resting heart rate, sleep efficiency, strain, respiratory rate — all required for physiological correlation. Mental log (via /mental): raw text or voice transcript, required for LLM sentiment and stress extraction.
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?Mental log entries are optional by design — the athlete writes only when they have something to say, with no format restrictions and no daily requirement. This preserves the signal quality of mental data: entries reflect genuine psychological states rather than forced daily check-ins. Impact on AI output: when mental log entries exist for a given date, LLM extracts stress level (1–10), motivation (1–10), sentiment, and key topics, which are then available for correlation with physical performance metrics. On days without a mental log entry, those fields are simply absent from the dataset and excluded from relevant pattern calculations.
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)Output is considered good when it meets the following measurable standards: pattern ρ value matches ground truth within ±0.05, p-value is below 0.05 for all surfaced correlations, RAG returns at least one topically relevant research reference per pattern, LLM response contains no metrics or claims absent from the input dataset, /ask response is delivered within 10 seconds, /mental processing is completed within 60 seconds, and output format follows the defined template with correct delta indicators and statistical qualifiers.
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?Output requires human judgment in two areas. First, pattern relevance: the athlete must decide whether a surfaced correlation makes intuitive sense given their training context — this is why the verification step (✓/✗) exists and cannot be automated. Second, tone and clarity: a sports scientist reviewing the output should find it precise and actionable, not vague or overly technical. These criteria are assessed during athlete testing sessions and coach review of monthly PDF reports.
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.Persona: a performance analytics assistant for a competitive baseball pitcher — precise, data-first, no motivational filler. System instruction: "You are a performance analytics assistant for a competitive baseball pitcher. Your role is to process raw athlete data and surface statistically grounded insights. Never speculate beyond the data. When referencing patterns, always cite the correlation coefficient (ρ), p-value, and supporting research. If data is insufficient, say so explicitly." Input structure: data is passed as labeled blocks — daily log fields, Whoop biometrics, and mental note transcripts are separated with clear tags. Constraints: never produce output that references metrics absent from the input, always include statistical qualifiers, never use hedging language. Variations to test: prompt with and without few-shot examples of good pattern descriptions, prompt with explicit instruction to flag low-confidence outputs.
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?Version 1 is the starting point — no iterations have been made yet as the bot is in pre-development. Prompt evolution will be tracked in a dedicated log file recording the version number, change made, reason for change, and observed impact on output quality. Expected iteration triggers: athlete finds pattern descriptions unclear during testing, RAG responses return irrelevant references, LLM produces overly technical language that requires simplification for non-expert readers such as coaches.
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?3 data sources are joined by date into a unified dataset before any analysis: daily logs from the Telegram bot (structured JSONB fields), Whoop biometrics fetched nightly via API, and mental log entries processed by LLM. Data preparation steps: date alignment across all three sources, addition of lag-1 columns (hrv_prev, recovery_prev, sleep_eff_prev) to enable next-day correlation analysis, and exclusion of rows with missing values for the specific metric pair being tested. RAG implementation: 6 peer-reviewed research articles are parsed, cleaned, and split into chunks of approximately 500 tokens each. Each chunk is embedded using OpenAI text-embedding-3-small and stored as a vector in PostgreSQL with pgvector extension. At query time, the pattern description is embedded using the same model and the top-3 most semantically similar chunks are retrieved. Claude then uses those chunks as context to return a science-backed assessment: supported, contradicted, or no_evidence.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.Input: 180 days of synthetic data for pitcher Jake Morris — daily logs with pitch count, arm fatigue, command rating, Whoop biometrics including HRV and recovery score, and mental log entries with stress and motivation levels. Expected output: Pattern Agent surfaces 3 statistically significant correlations — HRV previous day predicts pitch count (ρ = +0.86), recovery score previous day predicts pitch count (ρ = +0.63), HRV previous day predicts arm fatigue (ρ = −0.51). All three patterns return science status "supported" via RAG cross-check against 6 research articles.
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)Missing data: fewer than 30 training days logged — Agent returns "insufficient data" for training-specific pairs, no patterns surfaced. Ambiguous input: mental log voice note with background noise and incomplete sentence — LLM extracts available signal, flags low confidence, does not hallucinate missing context. Out-of-domain: athlete asks /ask about a metric not tracked by the system — bot responds "this metric is not available in your dataset" without attempting to generate an answer. Negative case: two unrelated metrics tested as a pair (resting heart rate vs. strikeouts) — Agent discards the pair (p > 0.05), nothing surfaced to athlete. Sparse Whoop data: API fetch fails for 2 consecutive nights — Agent skips those dates in correlation calculation, logs the gap, does not surface incomplete patterns.
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?The bot is currently in pre-development. Manual review of the Pattern Agent has been completed on 180 days of synthetic data. The agent correctly surfaced 3 out of 4 embedded patterns. The fourth pattern (stress level → nerves on game day) was not surfaced due to insufficient game day sample size — only 28 game days in the dataset, falling below the reliable threshold. This finding confirms the minimum data requirement of 50+ days per day type for reliable pattern detection.
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?Pattern Agent automated evaluation on synthetic dataset: 3 out of 4 embedded patterns detected (75% pass rate). All 3 surfaced patterns met statistical thresholds (|ρ| > 0.25, p < 0.05). All 3 patterns returned "supported" status via RAG cross-check. Correlation values matched embedded ground truth within ±0.05 for all 3 patterns. Full automated evaluation against real athlete data is pending bot development and data accumulation.
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?Two edge cases identified during Pattern Agent testing on synthetic data. First: insufficient game day sample — 28 game days produced unreliable correlations for game-specific metric pairs, confirming the 50-day minimum threshold. Second: missing Whoop data for 3 consecutive days due to simulated API gap — Agent correctly skipped those dates and excluded them from correlation calculations without surfacing incomplete patterns. Additional edge cases expected to emerge during real athlete testing: irregular logging behavior, extended rest periods, and mid-season injury gaps.
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?One prompt adjustment made based on manual review results: added explicit instruction to flag patterns with sample size below 50 days per day type as "insufficient data" rather than surfacing them with low confidence scores. This prevents misleading correlations from reaching the athlete. No further prompt iterations have been made — the system is in pre-development and full iteration cycles will begin after the bot is deployed and real athlete data accumulates.
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?Primary evaluation method is script-based: the Pattern Agent runs automatically on the full dataset and outputs results to patterns.json, which is then compared against the embedded ground truth programmatically. This allows fast, repeatable evaluation without manual review for each run. Human review is used as a secondary layer — the athlete verifies each surfaced pattern via ✓/✗, providing real-world signal on pattern relevance that no automated script can replicate. Scaling approach: as the athlete base grows beyond one pitcher, the evaluation script will run against each athlete's dataset independently, since patterns are individual and not averaged across users.
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?Pattern Agent runs every 3 days automatically — this is the core evaluation cycle. After each run, results are compared against previous patterns to detect new correlations and track confidence score changes on existing ones. Post-launch monitoring: evaluation script will re-run after every prompt update to confirm output quality is maintained. Monthly review of all verified patterns will be conducted to identify any patterns that have weakened or reversed direction as new data accumulates.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?The system is in pre-development — full infrastructure testing is pending bot deployment. Current readiness status: PostgreSQL schema with pgvector extension is designed and documented, API integrations for Claude, Whisper, OpenAI embeddings, and Whoop are selected and tested at script level, Railway is chosen as hosting platform with environment variables for all API keys. Pending before deploy: rate limit handling for Whoop API and Telegram bot, database backup and rollback strategy, monitoring setup for Pattern Agent and Notify Agent scheduled runs, error logging for failed API calls.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?No separate support, legal, or comms teams at this stage. Documentation complete: full project documentation is available in English covering system architecture, bot flow, pattern logic, data schema, and phase-by-phase progress. Athlete onboarding is handled directly by the team for the MVP pilot with one pitcher. Legal and privacy considerations for Whoop API data handling and athlete data storage are pending review before public launch.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?Pilot launch with one athlete — Simon, a real baseball pitcher. Closed pilot runs for a full season (approximately 6 months) to accumulate sufficient data for reliable pattern detection. Access is invite-only at this stage: no public onboarding, no self-signup. The team monitors data quality, logging consistency, and pattern relevance directly with the athlete. After pilot validation, rollout expands to a small cohort of 5–10 pitchers to test multi-athlete data isolation and sport template flexibility before any broader release.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?Database is designed for multi-athlete scale from day one — each athlete's data is isolated by athlete_id, sport templates are stored as JSONB config rows meaning adding a new sport requires no code changes. Railway hosting scales horizontally with usage. Pattern Agent and background agents run on APScheduler — scheduling logic is athlete-agnostic and will process each athlete's dataset independently as the user base grows. Initial volume will be monitored via Railway logs and database query times. Scale triggers: if response latency exceeds defined thresholds or Pattern Agent runtime grows beyond 5 minutes per athlete, infrastructure will be reviewed and optimized.
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?For the MVP pilot the following assets will be prepared: an interactive browser-based demo (single HTML file) showing the full athlete journey from data logging to pattern discovery, a one-page product overview explaining the core value proposition for data-driven pitchers and coaches, and a sample monthly PDF report showing what the athlete receives with real pattern examples and scientific references. Post-pilot assets will include an onboarding guide for new athletes covering Telegram bot setup and Whoop integration, and an FAQ covering data privacy, pattern reliability, and minimum data requirements.
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?Meeting Minutes — chronological record of each sync: what was discussed, raised, and decided. Captures context and open questions, not just conclusions. Decision Log (core) — append-only, deduplicated list of current locked decisions: date, decision, rationale, status (locked / revisit). Answers "what's true now" at a glance and prevents re-litigating settled choices. Every decision in the minutes is promoted to the log before the meeting ends; minutes and log cross-reference by date. Single source of truth — Google Drive holds the PRD, data, prototype links, roadmap, specs, minutes, and log. Rule: nothing important lives only in chat; decisions and changes are promoted into Drive. Async channel (Telegram/Slack) — day-to-day updates, blockers, quick back-and-forth. Ephemeral by design. Task board — shared To do / Doing / Done so each person's current work is visible. Weekly 30-min sync — fixed agenda (shipped / blocked / deciding); minutes taken, decisions logged before the call ends. Milestone recaps — at each phase gate: goal, outcome, metrics, what changed, next milestone.
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?Athlete data is stored in a private PostgreSQL database on Railway — accessible only by the development team. All API keys and credentials are stored as environment variables, never in code. Whoop biometric data is fetched via official Whoop API under the athlete's own account authorization. No athlete data is shared with third parties. Data sent to Claude API and OpenAI for LLM processing and embeddings is subject to Anthropic and OpenAI data processing terms — no training on user data under current API agreements. Full data privacy policy and athlete consent agreement are pending legal review before public launch.
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?Content moderation is not applicable — PitchIQ processes structured athletic data and free-form personal notes from a single athlete, not user-generated public content. No audit process is in place at MVP stage — this will be established before scaling beyond the pilot. Regulatory considerations: Whoop biometric data may fall under health data regulations depending on jurisdiction — legal review is required before expanding beyond the US pilot. GDPR compliance will be required if the product expands to European markets. At MVP stage with one consenting athlete, the team operates under direct agreement with the pilot participant.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?User metrics: athlete logs at least 5 days per week consistently over the pilot period, at least 3 verified patterns accumulated within the first 90 days, athlete reports that patterns are actionable and relevant during weekly check-ins, and monthly PDF report is shared with coach at least once during the pilot. Business metrics: successful completion of the full pilot season with one athlete, at least one pattern confirmed by both statistical analysis and athlete real-world experience, and a validated sport template structure that can onboard a second sport without code changes.
AI MetricsHow will you measure AI performance and accuracy?Pattern detection rate: agent surfaces at least 3 statistically significant patterns per 180 days of data. Correlation accuracy: surfaced ρ values match ground truth within ±0.05 on synthetic dataset. RAG precision: at least 80% of returned science references are topically relevant to the pattern being validated. Hallucination rate: zero instances of LLM referencing metrics or correlations absent from the input dataset. Latency: /ask responses delivered within 10 seconds, /mental LLM processing completed within 60 seconds. Athlete verification rate: at least 70% of surfaced patterns are verified (✓) rather than rejected (✗) — indicating the agent is surfacing relevant correlations, not noise.
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?At MVP pilot stage with one athlete, support is handled directly by the development team via a dedicated Telegram chat. No ticketing system or formal support infrastructure is needed at this stage.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?Feedback is gathered through three channels: weekly check-in conversations with the pilot athlete, athlete's ✓/✗ pattern verification decisions which signal whether the agent is surfacing relevant correlations, and direct observation of logging consistency which indicates whether the bot UX is frictionless. Critical issues such as agent failures, missing data, or incorrect pattern outputs are addressed immediately by the development team. Non-critical feedback is logged and reviewed at the end of each Pattern Agent cycle every 3 days. Product decisions based on feedback are documented in the project progress log.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?Railway platform provides baseline infrastructure logging — server uptime, response times, and error rates are monitored automatically. Application-level logging covers: Pattern Agent run status (success / failure / patterns found), Whoop API fetch results each night (success / skipped / failed), Telegram bot command usage and error responses, and LLM API call failures with error codes. Any Pattern Agent run that returns zero patterns or fails silently triggers a manual review by the development team. Mental log processing failures are logged separately to track Whisper and Claude API reliability.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?Learnings are collected through three continuous loops. First, every Pattern Agent cycle every 3 days produces new data on correlation strength and pattern evolution — reviewed jointly by the team. Second, athlete verification decisions (✓/✗) are tracked over time to identify whether pattern quality is improving as data accumulates. Third, prompt iterations are logged with version number, change, reason, and observed impact — reviewed monthly. The predefined metric pair list will expand over time as domain knowledge deepens and new correlations become worth testing. Adding a new sport beyond baseball requires only a new sport template row — no structural changes to the system.
Download the .xlsx ↓