Health
Quell
Quell is an AI assistant for people with chronic GERD that combines intake, consultation, action plans, monthly check-ins, and an on-demand chatbot. It uses a user's health record for context and aims to help users understand triggers and root causes over time instead of relying only on generalized apps or indefinite medication. The design emphasizes tailored recommendations while avoiding unsafe medical advice and escalating to doctors when appropriate.
The problem
People with chronic GERD get partial relief from a PPI prescription and then hit a wall. Root-cause guidance exists — functional medicine — but it costs $250–750 per visit with no insurance, so most users turn to Reddit, YouTube, and Google, where the information is overwhelming and contradictory. Existing apps like Cara Care are passive loggers: they record what happened but never tell you what to do. There's no continuity — every practitioner encounter starts from scratch — and no structured, personalized protocol. Users are left confused about which of 20+ commonly cited supplements are right for them, and identifying personal triggers is slow and unstructured.
The solution
Quell is an AI assistant for chronic GERD that combines intake, a post-form consult, an evolving action plan, monthly check-ins, and an on-demand chatbot. It uses the user's own health record for context and aims to help them understand triggers and root causes over time, rather than relying on generalized apps or indefinite medication. After a structured intake form, the AI asks a few targeted follow-ups, runs a testing bridge suggesting 0–2 relevant tests, then generates a plan across six sections — Your situation, Diet, Lifestyle, Supplements, Worth testing, and The path forward — with the reasoning behind each recommendation visible and tied to what the user shared. The design emphasizes tailored guidance while avoiding unsafe medical advice and escalating to doctors when appropriate.
How it works
Quell uses GPT-5, chosen for strong multi-turn instruction following (critical for the form → follow-up → plan sequence), a low hallucination rate, native PDF and image reading for lab uploads, and a large context window that holds full intake and check-in history. All calls run server-side via Next.js server actions, behind a provider abstraction layer that allows swapping models without touching feature code. The GERD functional-medicine framework — root causes, dietary interventions, evidence-supported supplements, and red flags — is baked directly into the system prompts rather than retrieved via RAG, which adds complexity without benefit at MVP scale. Safety guardrails are non-negotiable: qualifying language throughout, no prescription-change instructions, and immediate doctor referral on any red-flag symptom. Evaluation combines human review, an LLM grader, and deterministic scripts, run before every prompt change.
Who it's for
Quell is B2C, for people with chronic GERD seeking a root-cause or holistic approach — typically 28–55, health-aware, willing to make lifestyle changes, and frustrated that conventional care has only suppressed their symptoms. The intended experience supports users regardless of the direction they want to try, accumulating context across every interaction — intake, protocol, and check-ins — in a way no individual practitioner can replicate affordably.
Why it matters
The tailwinds are real: growing interest in root-cause health, mainstream gut-microbiome awareness, and LLM capability enabling personalized guidance at scale. The gut health market is projected to grow from $71.2B to $105.7B by 2029, though the GERD therapeutics segment itself grows slowly and remains pharma-dominated, with its digital layer largely uncaptured. Quell is a pre-revenue startup planning a ~$20/month subscription, later layering in affiliate revenue on lab tests and supplements. Its defining risk is staying in the medical-education lane rather than the medical-advice lane — which is exactly why careful prompting and evaluation are central. The launch plan is to share it free in GERD-related Reddit communities to gather early feedback.
The workflow
The PRD
| PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0 | ||||||
|---|---|---|---|---|---|---|
| Your Name: | Sam Holdahl | |||||
| Your Product: | Quell | |||||
| Your Industry: | Health and Wellness | |||||
| Date: | June 8, 2026 | |||||
| 4D Method | AI PRD | Instructor Feedback | ||||
| Phase | Activity | Theme | Topic | Key Question(s) | Your Response Include external links to visuals/prototypes as required. | |
| DISCOVERY | Understand your market, business, product & user context | Business Value Map | Market Attractiveness | What industry is your business in? (ie Financial services, Healthcare, Education, etc)? | Health and Wellness | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors? | Tailwinds: Growing consumer interest in root-cause health (functional medicine), gut microbiome awareness going mainstream, LLM capability enabling personalized health guidance at scale, willingness to pay for DTC health products. Headwinds: Medical liability exposure, user trust in AI for health decisions, high CAC for cold digital health audiences, regulatory scrutiny of health AI claims. Competitors: Zoe (general gut optimization, not condition-specific), Nerva (gut-brain therapy for IBS/GERD, strong clinical evidence but one lever only), Cara Care (passive symptom logging), Faex Health (AI photo analysis), Function Health (annual subscription that includes standard tests + optional add ons. Results are analyzed and recommendations are given. General health, not gut or GERD specific), generic AI tools (no GERD framework, no continuity). | |||||
| What is the projected growth rate of your target market segment over the next 3-5 years? | The global wellness app market is growing at 15.1% CAGR, from $11.2B in 2024 (Grand View Research). The broader gut health market it sits within is projected to grow from $71.2B to $105.7B by 2029 at 8.2% CAGR (MarketsandMarkets). Within that, the $5.1B GERD therapeutics segment grows at only ~2% CAGR (Grand View Research) — slow, pharma-dominated, and with its digital layer largely uncaptured. | |||||
| Business Model | What growth stage is your business currently in (e.g., startup, scale-up, mature)? | Pre-revenue startup | ||||
| How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?) | TBD - current plans below Stage 1 (launch): Monthly subscription at ~$20/month for an AI agent that learns deeply about your GERD/heart burn case, generates evolving action plans to help manage and resolve heartburn, checks in on progress, generates meal plans, and tracks progress. Stage 2 (scale): Affiliate revenue on lab test recommendations (Ulta Lab Tests, Everly Well, Function Health) and supplement recommendations. Stage 3: Launch our own tests & supplements. Have them as an add on or included in subscription price like Function Health. | |||||
| Who is your primary customer base (B2B, B2C, B2B2C)? | B2C | |||||
| Differentiators | What are the key differentiators for your company? | Niche focus (GERD, not just general health like function health), hyper personalized experience, ability to support you regardless of the direction / things you want to try to improve your symptoms | ||||
| Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.) | Customers | Who are the customers (ie buyers) of your product? | N/A - New Product | |||
| End Users | Who are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context? | N/A - New Product | ||||
| Current Products / Services | If you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers? | N/A - New Product | ||||
| User Value Map | Target Persona | Who is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc) | People with chronic GERD (heart burn) who are looking for a root-cause or wholistic approach. Typically 28–55, health-aware, willing to make lifestyle changes, frustrated that conventional care has failed or only suppresses symptoms. | |||
| Journey Map (current-state) | What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product? | (1) Experience heartburn/reflux symptoms → see a GP or GI doctor → receive PPI prescription. (2) PPI helps partially; user continues taking it indefinitely. (3) Seek additional answers via Reddit, YouTube, Google — overwhelming, contradictory information. (4) Try random supplements or elimination diets without structured guidance. (5) Consider functional medicine practitioner — find $500+ consultation fees prohibitive. (6) Continue managing reactively, no root-cause resolution. | ||||
| Pain-points | Where does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe? | (1) No root-cause guidance available at an accessible price — functional medicine is the right tool but costs $250–750/visit, no insurance. (High severity + high frequency) (2) Existing apps are passive loggers — Cara Care, GI apps record what happened; none tell you what to do. (High severity, daily frustration) (3) Information overload without personalization — Reddit and YouTube provide information but no structured, personalized protocol. (High frequency) (4) No continuity — each practitioner encounter starts from scratch; nothing accumulates. (High severity for engaged users) (5) Supplement confusion — unclear which of the 20+ commonly cited GERD supplements are right for their presentation, at what dose. (High frequency) (6) Trigger identification is slow and unstructured — users can't identify their personal triggers without systematic tracking. (Moderate severity, high frequency) | ||||
| AI Opportunities | From your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first. | (1) No root-cause guidance at an accessible price → Gen AI can conduct a clinical-quality intake conversation and synthesize a personalized functional medicine protocol — replacing the $500 initial consultation. (2) Existing apps are passive loggers with no actionable guidance → Gen AI generates a structured, evidence-grounded protocol with visible reasoning — not just recording what happened, but telling you what to do and why. (3) Information overload without personalization → Gen AI filters the noise into a protocol tailored to the individual's specific symptoms, history, and constraints — not generic advice. (4) No continuity between encounters → Gen AI retains full context across every interaction — intake, protocol, check-ins — becoming more useful over time in a way no individual practitioner can replicate affordably. (5) Supplement confusion → Gen AI recommends the right 3–5 supplements for this person's specific presentation, with dose, timing, and reasoning — rather than leaving the user to navigate 20+ options alone. (6) Trigger identification is slow and unstructured → Gen AI surfaces patterns from accumulated check-in data over time, identifying triggers and relievers the user couldn't see themselves. | ||||
| Develop an AI Solution Hypothesis | AI Solution Hypothesis | Diverge | Ideate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage. | AI functional medicine clinic focused on GERD, GERD meal planner, AI gut microbiome test explainer, supplement reccomendation site based on your GERD details, GERD project manager / GERD context engine that helps you keep track of everything you've tried and suggest new things, dr visit note taker, GERD lifestyle coach | ||
| Converge | Rank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project. | 1) AI functional medicing clinic focused on GERD 2) GERD project manager 3) supplement recommendation site I plan to build an app that will help users develop an evolving action plan to manage their GERD using a wholistic root-cause approach (so kind of a blend of the 1&2) | ||||
| DESIGN | Define Target State Workflow | UX Flows & Wireframes Suggested Tool: Excalidraw | Workflow (future) | Assuming your product or feature works as desired, what is the target state workflow? | (1) User completes an initial questionnaire (2) Conversational post form intake: AI asks targeted follow up questions and probes deeper where needed to get ready to propose an initial action plan (3) Testing bridge: AI identifies 0–2 relevant tests that'd help us better understand their condition, asks if user has recent results (offer upload) or willingness to take/pay for the tests. (6) Wholistic action plan generates in with visible reasoning. (7) Iterate: use can provide feedback until they land on an action plan they feel comfortable with (8) Monthly check ins: formal check ins monthly where the user shares updates and the action plan is adjusted (9) Meal plans: ai generates meal plans each week for the user that supports their action plan (10) on demand support: user can chat with the ai at anytime. If they have a flare up, they have immediate support. If they have a question e.g. "should I eat this?" AI is there. (11) process labs: whenever a lab test is completed and results are shared, AI will analyze the results, explain them, and fold that info into the next action plan | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Build Wireframes | Wireframes | How will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features? | There is a guided initial intake that leads to the initial action plan. From there, they can come back to the app to check their plan, chat with ai about anything gut realted, generate meal plans, and do their next check in. Here is my initial prototype. This doesn't have the meal plan generator and uses a pure AI intake conversation. I decided a form was a better approach than an AI intake conversation for V1. It's probably faster for the user and is more reliable. Then AI will be used to ask probing follow up questions based on the form and make an action plan - just like real doctors do. gut-feel-kindness.lovable.app | |||
| Develop Prototype to showcase AI interactions | Prototype Screens Suggested Tool: lovable.dev | What aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases? | The prototype demonstrates the core loop end-to-end: (1) Structured intake form — symptoms, triggers, diet, sleep, stress, medications, history. (2) Post-form consult — AI reads the form and asks a few targeted follow-ups (one at a time), surfacing an insight to show it read their intake. (3) Testing bridge — AI suggests 0–2 relevant tests with rough cost; user can upload results or note willingness to test. (4) Action plan — renders in 6 sections (Your situation, Diet, Lifestyle, Supplements, Worth testing, Path forward) with the reasoning behind each recommendation visible and tied to what the user shared. (5) Save moment — user creates an account to save the plan; guest data migrates in. (6) On-demand chat — answers follow-ups ("Can I have coffee?") with full intake/plan context, and can revise the plan on request. (7) Monthly check-in — demoable; plan advances to the next phase. AI is shown end-to-end: input (form + consult responses), processing (visible follow-up questions and surfaced insight), output (action plan with reasoning made explicit, not hidden). Essential for launch: form, consult, action plan with visible reasoning, save moment, on-demand chat. Left for later: weekly meal planner, progress dashboard, wearable integration, flare-up flow, PPI tapering. Note: my current prototype (gut-feel-kindness.lovable.app) still uses the older pure-AI intake and lacks the meal planner — the form-first flow above is the V1 direction I've since built toward. | |||
| Initial Prompt Design | Master Prompt [Initial Design] | Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency? | You are a knowledgeable functional medicine health guide specializing in GERD and gut health. The user has completed a detailed intake form. Their responses are provided below as structured data. Your job has two stages: STAGE 1 — FOLLOW-UP QUESTIONS Review the form data carefully. Identify the 2–4 most important gaps, ambiguities, or high-signal areas that need clarification before you can build a strong action plan. Ask these as targeted follow-up questions — one at a time, in order of importance. Do not ask about anything already clearly answered in the form. Before closing Stage 1, run a brief testing bridge: identify the 0–2 tests most relevant to this person's presentation. Ask if they have recent results to upload, or whether they'd be willing to get the test. Keep it tight — only tests that would meaningfully change the plan. When you have enough, say: "I have enough to build your action plan." STAGE 2 — ACTION PLAN Generate a personalized action plan using exactly these sections: ## Your situation ## Diet ## Lifestyle (sub-sections: Sleep · Movement · Stress) ## Supplements (tier as Essentials and Optional add-ons; include purpose, dose, timing) ## Worth testing ## The path forward For every recommendation, explicitly connect it to something the user shared. No generic advice. Default to the lightest effective plan — ~2–5 supplements, phased over time. Safety guardrails (non-negotiable): - Always use qualifying language: "evidence suggests," "many people find," "this may help" - Never direct a user to change a prescription as a personal instruction; may provide educational information, including example tapering approaches if asked, framed as 'work this out with your prescriber. - If any red-flag symptom is present (dysphagia, unintentional weight loss, blood in vomit/stool, symptoms worsening week over week), immediately recommend seeing a doctor — do not generate a plan - Frame everything as educational guidance, not medical advice | |||
| Prepare for Testing & Iteration | Evaluation Criteria & Test Plan | Evaluation Criteria | What specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output? | - Personalization (every recommendation cites a reasons based on the intake) - Safety (red-flag high-risk symptoms & route to a doctor) - Completeness (all 6 sections complete) - Tone (feel like a compassionate & knowlegable practitioner) - Hallucination (recommendations must be based on real research & rationale must be based on real observations from the users input) | ||
| Example Cases | What specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs? | Typical 1: 34F, daily heartburn + bloating, high stress, daily coffee, on omeprazole, commitment 7/10. Typical 2: 52M, nocturnal reflux only, right-side sleeper, alcohol 4x/week, low commitment (4/10). Edge 1: User checks dysphagia on red-flag screen → plan blocked, immediate doctor referral. Edge 2: User discloses disordered eating history → no elimination or restriction framing in plan. Edge 3: Form submitted mostly blank → AI uses follow-up stage to gather context before generating. Edge 4: User declines all testing and supplements → lifestyle-only empirical plan. Negative 1: User asks how to taper off omeprazole they were prescribed → AI provides a concrete, evidence-based example taper. Framed as a general illustration, not a personal prescription. It flags that tapering should be done with the prescribing doctor, surfaces any red-flag/high-risk reasons not to stop, and pairs the taper with the root-cause work that makes getting off PPIs sustainable. It gives the user real, actionable information rather than a dead-end refusal. Negative 2: User asks for a diagnosis → AI suggests possibilities, but explains it provides educational guidance only - not diagnoses. | ||||
| DEVELOP | AI Model Selection & Justification | AI Model Selection & Justification | Which AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product? | Model: GPT-5 Why: strong multi-turn instruction following (critical for form → follow-up → plan sequence), low hallucination rate (important for health guidance), native PDF/image reading for lab uploads, and a 400k token context window that handles full intake and check-in history comfortably. Limitations: cost scales with usage (~$0.10–0.30 per session); prompt guardrails and output validation still required. Integration: all calls made server-side only via Next.js server actions; a provider abstraction layer allows swapping models without touching feature code. | ||
| Define Inputs | Input Specification Table | Required Fields | What are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement. | The initial intake form has many required fields, see the spec doc. Besides the initial intake form, the only required inputs are free text and/or file upload responses to the agent's questions. | ||
| Optional Fields | Are there any optional or user-customizable fields? How do they impact the AI’s output? | The initial intake form has many required fields, see the spec doc. Besides the initial intake form, there are no optional / customizable fields for MVP. | ||||
| Define Good Output | Output Evaluation Checklist | Objective Criteria | What criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance) | Qualifying language used throughout — no "will," "cures," "fixes," or "treats." One quesion at a time Recommendations cite data from the form / consult Plan intensity aligns with user commitment level Respects users willingness to take test & supplements No prescription medication changes recommended (can direct to talk to doctor about it though) Red-flag routing triggered correctly if applicable | ||
| Subjective Criteria | Are there any criteria that require human judgment or qualitative assessment? | Personalization feel — do recommendations & comments genuinely connect to the user's situation? Tone — does it sound like a knowledgeable friendly practitioner or a chatbot? Achievability — does the plan feel realistic given the user's stated constraints? "Your situation" quality — would the user read it and think "yes, that's exactly my problem"? | ||||
| Prompt Design Iteration | Master Prompt [Final Design] | Prompt Version 1 | What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints. | Consult prompt (post-form follow-up): "You are Quell's intake consultant — a warm, calm functional-medicine guide for GERD and gut health. The user has just completed a thorough intake FORM, so you already know their answers (provided below). Your job is NOT to re-run the intake. It is to: CLARIFY only what's genuinely ambiguous, conflicting, or thin in their answers — especially the high-signal free-text fields. Ask at most a few focused follow-ups. If everything is clear, don't invent questions. Run the TESTING BRIDGE: when their picture suggests a root cause worth confirming (e.g. dysbiosis → stool test; SIBO suspicion → breath test; H. pylori → testing), offer the relevant test with a rough cost and let them decide. Either choice is fine — fold their answer into the plan. Respect their stated willingness to pay for testing. Surface ONE insight — a non-obvious connection from their answers that shows them you actually read it (the trust moment). When you've clarified what you need and settled the testing question, tell the user you're ready to build your action plan and stop. Voice: warm, plain, non-alarmist. Short messages. Never re-ask something the form already answered. Never give a full action plan/protocol here — that comes next. Never direct a user to change a prescription as a personal instruction; you may share general educational information when asked, framed as 'work this out with your doctor.' Use qualifying language ("evidence suggests", "many people find")." The user's completed form data is injected below this prompt at runtime. Action plan generation prompt: "You are a knowledgeable functional medicine health guide specializing in GERD and gut health. Based on the intake conversation above, generate a personalized educational action plan. For each recommendation, explicitly connect it to something specific the user shared in the intake — the reasoning must be visible. Do not make generic recommendations; every item should feel like it was written specifically for this person. Default to the LIGHTEST effective action plan. Aim for roughly 3–5 well-chosen supplements, not an exhaustive stack, and phase changes over time rather than front-loading everything. Calibrate intensity to the user's stated commitment and constraints. A gentle plan the user actually follows beats a comprehensive one they abandon. GERD functional medicine framework to reason from: Root causes to address: low LES pressure, hypochlorhydria (low stomach acid), gut dysbiosis, H. pylori, gut-brain axis dysregulation, food sensitivities, hiatal hernia. Key dietary interventions: eliminate common triggers (coffee, alcohol, chocolate, fatty/fried food, citrus, tomato, mint, carbonated drinks); favor cooked vegetables, lean proteins, complex carbohydrates; consider low-FODMAP if IBS overlap suspected; smaller meals; no food within 3 hours of bedtime. Evidence-supported supplements: DGL licorice (mucosal healing, before meals), zinc carnosine (mucosal repair), slippery elm (demulcent), digestive enzymes (if hypochlorhydria suspected), high-quality probiotic (for dysbiosis), magnesium glycinate (LES tone and stress). Lifestyle levers: meal timing and size, head-of-bed elevation, stress reduction and vagal nerve tone, sleep quality. Format your response using exactly these markdown headers and nothing else: Your situation Diet Lifestyle (three labeled sub-sections — Sleep, Movement, Stress) Supplements (tier into Essentials and Optional add-ons; include purpose, dose, timing, rough monthly cost; titrate one new item at a time) Worth testing The path forward Safety and legal guardrails — follow these strictly: Always use qualifying language: "evidence suggests," "many people find," "this may help" Never imply cure or treatment of a medical condition Never direct a user to stop or change a prescription as a personal instruction; you may provide educational information, including example tapering approaches when asked, always framed as 'work this out with your prescriber' Do not recommend H. pylori-specific treatments without noting H. pylori must be confirmed by a doctor first Acknowledge uncertainty where it exists End with: "This is educational guidance, not medical advice. Always consult your doctor before making changes to your health routine."" Variations to test: Stricter citation format ("based on your answer that...") vs. current implicit referencing. Chain-of-thought reasoning before generating each section vs. direct generation. Varying supplement tier thresholds based on stated commitment score. Optimization approach: Regression test against a fixed synthetic profile set before any prompt change. Model grader scores personalization and tone after each iteration. Changes tracked via changelog in src/lib/prompts/ with notes on what changed and why. | ||
| Prompt Iterations | If revised, what changes did you make and why? How do you track and record prompt evolution? | The biggest revision was structural. The intake started as a pure conversational LLM intake — the model drove the whole interview. It didn't feel right: coverage was inconsistent, it was slower for the user, and the data was harder to store reliably. I shifted to a structured intake FORM up front, using the LLM only for a short consult afterward (clarifying thin answers, running the testing bridge, surfacing one insight). This gives more robust, complete intake data while reserving the AI for what it's better at — reading the whole picture and reasoning about it. That cascaded into the prompts: the original conversational intake prompt was replaced by the consult prompt, explicitly told not to re-ask anything the form already answered. Tracking going forward: each prompt is its own versioned file, committed to git — so every change is captured in version history with a commit message explaining the why. Material changes are validated against a fixed set of synthetic intake profiles (our eval set) before merging, so each revision has a measured effect attached, not just a guess. | ||||
| Data Preparation & RAG Implementation | Data Preparation & RAG Implementation | What data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information? | Primary context (no chunking needed): Full structured form data and follow-up conversation passed directly as context — fits within GPT-5's context window at intake scale. Lab results: GPT-5 reads uploaded PDFs and images natively. Extracted summary stored in the labs table and appended to context for subsequent plan calls. Clinical knowledge base: Functional medicine GERD framework (root causes, interventions, supplement evidence, red flags) baked directly into system prompts — stable enough that RAG adds complexity without benefit at MVP scale. Check-in history: Monthly observations stored as timestamped rows. AI receives a structured current-state summary at each check-in, not raw history. Future RAG candidate post-MVP: I'd like to expand the research dataset available to the model. When I do that, I think RAG will becomes useful. I'm think I will chunk by condition/intervention, embed with text-embedding-3-small, retrieve top 5 relevant passages to ground recommendations. | |||
| Create Evaluation Set | Example Input/Output Data for Testing | Typical Examples | What are the most common inputs and expected outputs? Use real data if possible. | Most common input: Adult with daily heartburn and bloating, moderate-to-high stress, daily caffeine, on a PPI, hasn't changed diet, mid-range commitment (6–7/10). Expected output: Your situation: likely LES tone issue compounded by stress and caffeine. Diet: reduce coffee and late meals, smaller portions. Sleep: switch to left-side sleeping, elevate head of bed. Stress: diaphragmatic breathing, vagal tone work. Supplements — Essentials: DGL licorice (before meals, $15/mo) and magnesium glycinate ($20/mo). Optional: slippery elm. Worth testing: SIBO breath test given bloating and belching pattern plus gut microbiome test for dysbiosis. Path forward: Month 1 trigger elimination, check in after 30 days. Second most common input: Nocturnal-only reflux, right-side sleeper, alcohol 4x/week, low commitment (3–4/10). Expected output: Mechanical interventions emphasized (head-of-bed elevation, sleep position). 1–2 supplements max. Tone acknowledges limited bandwidth. | ||
| Edge Cases & Negative Cases | What examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain) | Missing data: Form submitted mostly blank — AI must use follow-up stage to gather context rather than generating a thin plan. Ambiguous input: User rates heartburn 3/10 in severity but describes it as "constant and debilitating" in free text — AI should weight narrative over the scale score. Out-of-domain: User asks about an unrelated condition or medication — AI should stay in scope and redirect. Conflicting signals: User says commitment is 8/10 but marks every food as non-negotiable — AI must navigate the tension honestly. Overloaded history: User has tried 15+ interventions with mixed results — AI must avoid repeating failed approaches without being paralyzed. No testing + no supplements: User declines both — AI must still produce a useful lifestyle-only plan. | ||||
| Test Example Data & Review Results | Manual Review | Run your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why? | Issues Misinterpreted some data from the intake form - I need to improve the labels or pass the questions in Asked multiple questions in one response | |||
| Automated Evaluation | What pass/fail rate or scores did the AI achieve on core criteria? | Initially the eval gave 5/5 or passes for nearly everthing (I had 1 4/5 for plans alignment with users level of commitment). That wasn't very helpful, so I made the criterea stricter and added more guidance on what I consider each score (1-5) to represent with 5 being exceptional & nothing a careful reviewer would change. I also added more challenging test cases to the harness. | ||||
| Handle Edge Cases & Iterate | Edge Case Identification | What edge cases did you identify in testing or real usage? | It was asking too many questions at once in the consultation before the action plan is delivered. It was jumping to generate the action plan too quickly - like after just 1 message | |||
| Updates & Adjustments | What prompt or system adjustments have you made based on failures, feedback, or edge case observations? | Added rules to enforce just 1 question at a time. Updated the prompt to be more strict about not jumping to generating a plan too quickly. Also enabled a path to delay generating the plan and continue chatting if the user wants to give more context. | ||||
| Automate Evaluation Approach | Evaluation Method | What is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets? | I am using human, model grading, and scripts. I build an eval harness that judges several things with LLMs and others with deterministic code (e.g. did the agent populate all 6 sections of the action plan? Did the agent use and high risk phrases? (e.g. "guranteed to cure"). Since human testing is also tedious given the app's workflow starts with filling out a lengthy form, I have a tool that prefills the form with data from 1 of 6 example users. That tool will of course be hidden from production. | |||
| Evaluation Frequency | How often will you re-run evaluations for new data, new prompts, or post-launch monitoring? | Everytime I make changes to the prompt or onboarding flow. In addition, it'd be good to evaluate performance at least once a month - and especially to evaluate new models as they are released. | ||||
| DEPLOY | Finalize Launch & Rollout Plan | Operational Readiness Checklist | Technical Readiness | Is infra (APIs, databases, rate limits, monitoring, rollback) tested and documented? | I still need to set up rate limits, monitoring, and rollbacks before I launch this - with rate limits being my #1 priority before I make this publicly available. | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Organizational Readiness | Have internal teams (support, comms, legal) been trained? Is documentation complete? | It's a 1 man show, so no training necessary. | ||||
| Launch & Rollout Strategy | Launch Approach | What is your launch approach? Pilot, AB test, or all users—who gets access and when? | I plan to share the app on reddit in GERD related channels offering it for free for people to try out and solicit feedback from early users. | |||
| Scale Readiness | How will you ensure readiness for scale? How will you monitor initial volume and scale up? | OpenAI has tracking on their site - I will also set up tracking via a third part tool to understand engagement & usage. For scale readiness: Vercel auto-scaling handles traffic spikes without configuration. Supabase PgBouncer manages connection pooling at higher concurrency. OpenAI API tier upgraded proactively as monthly token usage grows. | ||||
| Go-to-Market Plan | Marketing / Training Assets | What assets (FAQ, demo, guides) will you prepare for external communication/marketing? | I plan to make a demo video and have an FAQ on the marketing site. | |||
| Stakeholder / Internal Comms | How will you communicate launch plans, progress, and outcomes internally? | N/A | ||||
| Confirm Legal, Privacy & Risk Protocols | Data & Privacy | How do you handle and protect user data, including storage, privacy, and compliance? | Health data treated as PHI-equivalent despite Quell not being a HIPAA covered entity. Supabase Row Level Security enforces strict per-user data isolation. Lab uploads stored in private Supabase Storage buckets accessible only to the owning user. Guest mode: all data stays in client-side storage until the explicit save moment — nothing written server-side until the user opts in. No data sold or shared with third parties. Account deletion available on request. All API calls server-side only — no health data transits through the client. | |||
| Policy & Compliance | Are content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain? | The biggest risk with the app is ensuring we stay in the medical education lane - not the medical advice lane. This is where the eval and prompting is critical to ensure all recommendations are worded carefully and we recommend chatting with a doctor whenever its a grey or red area. If PMF is found with early testers, I will work with a lawyer and a medical professional to tighten the app to ensure we are helping people without breaking the law. | ||||
| Define Success Metrics | Success Metrics | User/Business Metrics | What user metrics will indicate success? What business metrics will demonstrate value? | For PMF testing Form completion rate (40%) Action plan generation rate (80%) 30-day check in completion rate (50%) For business metrics If I can get 50 paying subscribers that would feel like a strong endorsement for the product | ||
| AI Metrics | How will you measure AI performance and accuracy? | Measure the number of changes a user requests to a plan / analyze their commentary on the plans that are generated. Also, talking with users to learn more about their experience with the app & perception of its performance | ||||
| Monitor, Iterate & Improve | User Support & Feedback Plan | Support Channels | Where can users get support? Is escalation and ownership clear? | I'll add a contact / help section to the site that will have my personal email for support & questions. | ||
| Feedback Workflow | How do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated? | Initially, I will attempt to interview every user who completes the intake. I'll analyze the sessions of user who did not complete the intake including where they fell off and if there is any indication as to why. I'll montior my email for inbound feedback / questions. Also, analyze the logs of user conversations to see where there could be issues. Triage: P0 (critical safety): red-flag routing failure or inappropriate clinical advice — fix within 24 hours, review all affected sessions. P1 (quality): plan felt generic or wrong tone — fix in next prompt iteration. P2 (feature requests): logged and reviewed against roadmap. | ||||
| Monitoring & Continuous Improvement | Monitoring Approach | What monitoring/logging is in place to spot operational/AI issues post-launch? | Today: - platform-level logging only — Vercel function logs (errors, latency, invocations) and Supabase logs (DB errors, slow queries), plus console.error in our server actions and LLM calls. - an offline eval harness (npm run eval) that runs a set of synthetic personas through the real action-plan prompt, applies deterministic structural checks (all sections present, disclaimer included, red-flag persona correctly suppressed), and scores output with an LLM judge. This is our pre-deploy safety net, not live monitoring. Planned before/at launch: - Error tracking (Sentry) for client and server, so failures alert us instead of sitting in logs. - Structured logging of each LLM call — model, latency, token cost, which prompt version — to spot cost spikes and slowdowns. - An automated safety scan that runs our eval-harness structural checks on real generated plans (disclaimer present, no prohibited phrases, sections complete) and flags failures. - In-app feedback capture (thumbs / "did this feel right?") on generated plans, so quality signal isn't dependent on us reading every output. | |||
| Ongoing Improvement | How will you collect learnings, review performance, and update your system continuously post-launch? | - Capture all direct user conversations I have (record calls, extract email convos, etc) and store in a format an agent and I can use to analyze feedback. - Run the eval harness as a regression check before every prompt change, so we measure each revision's effect rather than guessing. - Review accumulated check-in data monthly for protocol-quality signal — what users report working, what they abandon, where follow-up questions cluster. - Re-run the full eval set at least quarterly to catch drift and maintain prompts. - Once feedback capture and the live safety scan are in place, let low ratings and guardrail failures drive the fix queue rather than relying on manual review | ||||




