← All capstone projects

Health

BioBuddy

Built by Rodrigo Tello and Ansar Cohort 9 Digital health / personal health intelligence

BioBuddy is an AI-powered personal health intelligence platform that unifies blood test reports, Apple Health exports, and nutrition logs into one queryable knowledge base. The team focuses on cross-source health questions that existing tools do not answer well, such as relating biomarkers to sleep, energy, or calorie intake. The prototype normalizes messy source formats, aggregates nutrition data, and provides chat-based summaries and answers grounded in uploaded records.

The problem

People who take their health seriously already track it obsessively — Oura, WHOOP, Levels CGM, InsideTracker blood panels — yet the data lives in incompatible silos. A biohacker who wakes up flat spends twenty minutes hopping between four apps and a lab PDF, builds a rough hypothesis, and has no way to validate it or track whether the pattern repeats. Blood tests stay trapped in PDFs with no longitudinal tracking. Mood and journal entries never connect to objective biomarkers. Every existing tool forces you to look at one source at a time, and the whole process is reactive — you only investigate after you already feel the impact.

The solution

BioBuddy is a personal health intelligence platform that unifies blood test reports, Apple Health exports, and nutrition logs into one queryable knowledge base. It answers cross-source questions no single-metric tool can — "Did my HRV drop in weeks my ferritin was low?" or "How does my sleep look on high-calorie days?" — grounded in the user's own numbers with cited sources. The product combines a RAG-powered chat interface for deep reactive Q&A, a real-time health dashboard, and a proactive pattern-detection engine that surfaces concerning trends before the user notices. Insights are calibrated against both clinical reference ranges and the individual's personal baseline.

How it works

Three prompts govern the system: a clinical extraction prompt that pulls numeric biomarkers from any lab PDF into strict JSON, a pattern-detection prompt that only fires on 3+ consecutive data points, and a Q&A prompt that answers from the user's data first and never makes diagnoses. Under the hood, pdfplumber handles layout-aware table extraction, biomarkers are chunked per-marker, and text is embedded with nomic-embed-text via Ollama. Retrieval is hybrid — an LLM classifier routes each query to structured SQL, vector similarity search, or both. Inference runs on Groq's llama-3.1-8b-instant (128k context, sub-2s latency), with a FastAPI backend and Supabase PostgreSQL plus pgvector for storage.

Who it's for

BioBuddy is a B2C product for the biohacking and quantified-self community — people aged 25–45 who already spend on self-tracking hardware and lab work and have proven willingness to pay for health insights. The most revenue-generating segments are biohackers who order their own blood panels and want to correlate everything, and fitness-focused optimizers tracking recovery, macros, and performance. A third segment, health-anxious individuals managing diagnoses like thyroid or iron issues, wants to understand trends without a doctor visit. The common thread is a motivation to understand how inputs — nutrition, sleep, training — affect outputs like energy and biomarkers.

Why it matters

The biohacking and quantified-self market was valued at $24.81B in 2024 and is projected to reach $149.6B by 2029 — a ~32% CAGR — while AI in healthcare grows at roughly 48% CAGR. Wearable adoption, mainstream direct-to-consumer lab testing, and post-pandemic health awareness have made multi-source pattern detection feasible at the consumer level for the first time. The launch plan is a closed beta with recruited biohackers to validate Q&A quality and alert precision before opening paid tiers. AI targets are exacting — extraction accuracy above 95%, hallucination rate under 1% — and HIPAA and PIPEDA compliance are prerequisites before commercial launch.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Rodrigo Tello & Ansar Khan
Your Product:BioBuddy
Your Industry:HealthTech
Date:May 8th
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Consumer digital health, specifically the biohacking and quantified self market.Please leave this area blank. This space is for the Instructor to provide you with feedback.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?The primary headwinds are health data fragmentation across incompatible platforms, consumer privacy concerns around centralizing sensitive biometric data, and the regulatory grey zone around AI-generated health insights. On the tailwinds side, wearable and continuous monitoring adoption is accelerating, direct-to-consumer lab testing has made regular bloodwork mainstream, and post-pandemic health awareness has expanded the biohacker demographic. AI capability improvements have made multi-source pattern detection feasible at the consumer level for the first time. Competitors include InsideTracker (biomarker optimization), Heads Up Health (health data aggregation), Levels Health (metabolic health via CGM), Whoop and Oura (recovery and readiness), and Function Health (comprehensive lab testing). None offer a unified natural language Q&A interface synthesizing across all health domains simultaneously with proactive alerting.
What is the projected growth rate of your target market segment over the next 3-5 years?The biohacking and quantified self market was valued at $24.81B in 2024 and is projected to reach $149.6B by 2029, a CAGR of ~32.1% (Straits Research, 2025). The broader personal health analytics market shows a CAGR of 24.3% over the same period (MarketsandMarkets, 2025). AI in healthcare is growing at 48.1% CAGR through 2029 (MarketsandMarkets, 2024).
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Pre-launch, MVP stage.
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)BioBuddy operates on a subscription-tier model. A limited free trial allows users to explore the product with restricted data history and integrations. Paid tiers unlock full historical analysis, all data source integrations, proactive alerts, and unlimited Q&A. Pricing is structured as monthly and discounted annual plans.
Who is your primary customer base (B2B, B2C, B2B2C)?B2C.
DifferentiatorsWhat are the key differentiators for your company?BioBuddy provides a single natural language interface across all health domains simultaneously — blood biomarkers, nutrition, fitness, sleep, mood, and journal entries — in one unified knowledge base. Unlike single-metric platforms, it synthesizes patterns across domains to answer questions like 'why aren't my workouts as effective lately?' It combines proactive monitoring with reactive deep-dive Q&A, and calibrates insights against both clinical reference ranges and the individual user's personal baseline.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?Aimed at biohackers and fitness focused individuals who are already spending money on self-tracking tools — Oura rings ($300+), WHOOP subscriptions ($30/mo), Levels CGM ($200/mo), InsideTracker blood analysis ($500/year). These people have already proven willingness to pay for health insights and BioBuddy sits directly in that ecosystem and competes on depth of cross-source intelligence rather than hardware.
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?Biohackers - Obsessively track HRV, sleep stages, biomarkers; order their own blood panels (e.g. via Marek Health, Function Health); want to correlate everything Fitness focused optimizers - Athletes and gym-goers using Apple Health + nutrition tracking; care about recovery, macros, and performance metrics Health anxious individuals - People with a diagnosis (thyroid, iron deficiency, pre-diabetes) who get regular labs and want to understand trends without a doctor visit Biohackers & fitness focused optimizers are the most revenue-generating segments since they have strong existing spend on health tools — higher willingness to pay and lower price sensitivity. These users also have stronger word of mouth impact as they're frequently active in fitness communities and forums and will share info on tools that work for them. We assume the profile of these users will be - Age 25–45, tech-savvy, motivated by optimization not just maintenance - Core goal: understand how their inputs (nutrition, sleep, training) affect their outputs (energy, body composition, biomarkers)
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?BioBuddy is product-led. It ingests data from blood test PDFs, fitness trackers (e.g. Apple Health), diet & nutrition trackers (e.g. LoseIt, myfitnesspal, etc.), and a manual daily log then surfaces that data through three core features: RAG-powered chat interface The central differentiator. Users ask natural language questions that span all their data sources — "Did my HRV drop in weeks my ferritin was low?" or "How does my sleep look on high-calorie days?" — and get answers grounded in their actual numbers, with cited sources. This is the only tool that lets you query across blood, activity, sleep, and nutrition simultaneously. Every other tool forces you to look at one source at a time. Health dashboard A real-time overview of the metrics that matter most — resting HR, HRV, sleep, weight, calories — with time-series charts and adjustable windows. This addresses the "I don't even know where to start" problem: users get an at-a-glance picture of their health without needing to write a query or open three separate apps.
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)Health-invested individuals in the biohacker and quantified self community who already track multiple data streams (wearables, regular bloodwork, nutrition logs, sleep data) but have no unified way to synthesize that data into actionable intelligence. They are technically comfortable, motivated by performance optimization, and willing to pay for tools that give them a measurable, data-backed edge.
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?The user wakes up feeling flat and underperforms in their morning workout. They open Apple Health and look at sleep metrics. They open their nutrition app separately. They dig up their last lab report PDF and scan it manually. They check their Oura or Whoop app for recovery scores. After 20 minutes across four different apps, they have a rough hypothesis but no way to validate it or track whether the pattern repeats. The next time it happens, they start from scratch. Lab results sit in PDF form with no longitudinal tracking. Mood and journal entries live in separate tools with no connection to objective biomarkers.
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?1. No cross-source intelligence layer — data from blood tests, wearables, nutrition, and mood lives in completely separate silos (most severe). 2. Entirely reactive process — the user only investigates after already feeling the impact of a pattern. 3. Blood test data trapped in PDFs with no longitudinal tracking or trend detection. 4. Journal entries and mood data never linked to objective biomarkers.
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.1. Multi-source causal Q&A — synthesizes sleep, nutrition, blood, and fitness data to answer 'why do I feel X?' 2. Proactive pattern detection — monitors trends and surfaces concerning signals before the user notices. 3. Automated biomarker extraction — normalizes any lab PDF format regardless of lab or layout. 4. Personalized baseline learning — calibrates alerts and insights to the individual's own norms over time. 5. Journal and mood synthesis — links unstructured personal entries to objective data patterns.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.• Unified health knowledge base with natural language Q&A across all connected data sources • Proactive pattern detection engine that surfaces in-app banners for concerning trends • Automated lab result ingestion via email report parsing and lab portal APIs • Daily mood and energy logging with a quick in-app prompt • Journal entry ingestion from Notion or manual upload with LLM theme extraction • Weekly AI-generated health digest summarizing key trends and anomalies • Pre-lab appointment brief compiling the last 12 months of relevant data • Personalized baseline calibration that learns the user's normal ranges over time
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.Top 3: 1. Unified health knowledge base with multi-source RAG Q&A — highest user value, core intelligence layer 2. Proactive pattern detection and alert engine — highest differentiation, converts passive tool to active monitor 3. Automated multi-source data ingestion — required to keep the knowledge base continuously useful Selected: An intelligent personal health operating system combining automated multi-source data ingestion, continuous pattern monitoring with proactive in-app alerts, and deep reactive Q&A through a unified natural language interface.
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?Connected data sources sync automatically in the background (Apple Health, email lab reports, Notion journals). The user completes a daily 30-second mood and energy check-in in-app. The pattern detection engine runs continuously, comparing incoming data against clinical reference ranges and the user's personal baseline. When a concerning pattern is detected, an in-app banner surfaces the insight. The user can tap to ask a follow-up question, or open chat and ask 'why have my workouts felt harder this week?' The system retrieves data across all sources and synthesizes a grounded answer with cited values and dates.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?Five main views: • Home: Active insight banners, daily health snapshot (sleep score, HRV, energy, nutrition), quick-log card for mood/energy • Data Sources: Connected integrations with status indicators and last-sync timestamps • Dashboard: Historical metric charts, biomarker table with longitudinal comparison, correlation explorer, raw data tables • Chat: Persistent conversation window with citation panel, session management • Settings: Subscription management, notification preferences, data export
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?Current prototype demonstrates: PDF upload with biomarker table appearing in seconds, longitudinal chart comparing multiple test dates, and natural language Q&A with cited values and expandable source panel. Next iteration: triggered in-app alert banner and tap-through to contextual chat conversation. Essential for launch: PDF upload, Apple Health and nutrition ingestion, multi-source chat Q&A, biomarker dashboard. Post-launch: mood logging, journal ingestion, proactive alert engine, Notion integration, personalized baseline calibration.
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?Three prompts govern the system: EXTRACTION PROMPT — Tone: clinical and precise. The model acts as a medical data transcription tool. Instruction: extract all numeric biomarker results, map names to standardized snake_case via a BIOMARKER_ALIASES dictionary, parse all reference range notation styles, detect flags from explicit markers or implicit comparison, return strict JSON (no markdown, no explanation). PATTERN DETECTION PROMPT — Takes a structured trend snapshot across all sources and evaluates whether patterns cross an alert threshold against both clinical norms and the user's 30-day personal baseline. Only fires on 3+ consecutive data points. Q&A SYSTEM PROMPT — Tone: informative, not clinical. The model acts as an intelligent personal health analyst. Instruction: answer using the user's own data as the primary source, apply both personal baseline and clinical reference ranges, synthesize across all connected sources, supplement with established health guidelines where needed, distinguish personal data from general guidance, never make diagnoses or treatment recommendations.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?• Factual accuracy: all cited values match exactly what is in the database • Source attribution: every data claim references a specific date, source, and metric • Hallucination avoidance: no fabricated values, dates, or measurements in any response • Cross-source coherence: multi-source reasoning is logically consistent • Baseline awareness: answers account for the user's personal norms, not just population averages • Appropriate tone: informative without being alarmist or falsely reassuring • Alert precision: proactive alerts trigger only on genuine multi-point patterns, not single anomalies
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?Typical: 'Why haven't my workouts been as effective lately?' — pulls sleep, nutrition, HRV, and activity data; identifies contributors with cited values. Typical: 'What were my flagged blood test values?' — returns monocytes and transferrin saturation as flagged low with reference ranges. Edge: Question about an unconnected health domain — model states data is unavailable and suggests what to connect. Negative: 'Should I take magnesium?' — model declines the recommendation but shows relevant values. Edge: Journal entry with ambiguous mood language — model extracts themes conservatively. Negative: Pattern detection on fewer than 3 data points — no alert fires. Edge: Blood test PDF that is a scanned image with no extractable text — fails gracefully with clear error.
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?Three models are used: nomic-embed-text (768-dim, via Ollama): Handles all embeddings. Runs locally at zero cost, performs well on short biomedical text. Limitation: requires local Ollama instance; bottleneck at multi-tenant scale — migrate to hosted API before production. Groq llama-3.1-8b-instant: Handles extraction, classification, and Q&A. Sub-2s latency, native JSON output, 128k context window. Limitation: 8B parameters may be imprecise on complex tabular formats and nuanced multi-source synthesis. Q&A and pattern detection are candidates for llama-3.3-70b-versatile or Claude Haiku 4.5 at higher cost. Integration: Groq via official Python SDK; Ollama via local HTTP service.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.Blood test upload: PDF file (any format), auth token — multipart form / JWT Apple Health: XML export file, auth token — multipart form / JWT Nutrition (Lose It): CSV export, auth token — multipart form / JWT Daily mood log: mood rating (1–5), auth token — JSON / JWT Journal entry: text content, date, auth token — JSON / JWT Chat Q&A: natural language question string, auth token, session ID — JSON / JWT
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?Session ID (chat): enables multi-turn conversation memory. Without it, each message is treated independently. Date override (journal entry): for documents without a clear timestamp. Energy level rating 1–5 (mood log): companion to mood rating, improves pattern detection context. Free-text note (mood log): additional qualitative context used in journal synthesis and pattern analysis.
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)• Extracted numeric values match the source PDF exactly (no rounding or approximation) • All reference range notation styles parse correctly into numeric low and high bounds • Flag detection is consistent across all notation formats (explicit columns, symbols, implicit comparison) • Biomarker names resolve to the correct standardized snake_case name from the alias dictionary • Pattern detection fires only when a trend is present across at least 3 consecutive data points • Chat responses cite at least one specific value and date per data-based claim
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?Chat responses should be readable by a health-literate but non-clinical user — not over-simplified, not overly technical. The model should express uncertainty proportionally and not appear more confident than the data warrants. When synthesizing across multiple data sources, causal reasoning should be plausible and internally consistent. Proactive alerts should feel genuinely useful and actionable rather than anxiety-inducing or vague.
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.Extraction Prompt V1: 'You are a medical data extraction assistant. Extract all biomarker results from the blood test report text. Return ONLY valid JSON with: lab_name, test_date, patient_name, biomarkers (array of: name, standardized_name, value, unit, reference_range_low, reference_range_high, flag, reference_range_notes), notes. Extract every numeric result. Parse all reference range formats. Detect flags from explicit markers or by comparing value to range. Value must always be a number, never a string.' Q&A System Prompt V1: 'You are a personal health assistant. Answer primarily using the provided context from the user's health data. For general health guidelines and recommended daily values, you may draw on your training knowledge — but clearly distinguish between the user's own data and general guidance. Be specific with numbers and dates. Do not make diagnoses or treatment recommendations.'
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?Extraction prompt (5 iterations): 1. Simple extraction with basic JSON schema 2. BIOMARKER_ALIASES dictionary added for cross-lab name standardization 3. Comprehensive reference range format rules added (handles >, <, >=, <=, tiered text) 4. Explicit flag detection logic added for all notation styles 5. reference_range_notes field added to capture tiered thresholds and clinical guideline text Q&A prompt (2 iterations): 1. Started with 'answer ONLY from provided context' — model would say 'I cannot answer' for questions about well-established guidelines 2. Relaxed to allow established health guidelines as supplementary knowledge with clear attribution
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?Sources: Lab PDF reports (any format), Apple Health XML export, Lose It nutrition CSV, daily mood logs (structured), journal entries (Notion or manual upload). Cleaning: pdfplumber for layout-aware PDF table extraction; XML parsing for Apple Health; CSV normalization with duplicate deduplication for Lose It; LLM theme extraction for journal entries. Chunking: Per-biomarker chunks for blood tests (one chunk per biomarker with value, unit, reference range, flag, clinical notes, lab name, date); 500-character sentence-aware chunks with 50-char overlap for other documents; entry-level chunks with date metadata for journals and mood logs. Embedding: nomic-embed-text (768-dim) via Ollama. Document metadata (source type, date) prepended to chunk text before embedding. Retrieval: Hybrid — LLM classifier routes each query to structured SQL, vector similarity search (threshold 0.6, match count up to 20), or both. Biomarkers also stored as individual structured rows in health_metrics table for direct SQL retrieval.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.Input: 'Why haven't my workouts been as effective lately?' Expected: Hybrid retrieval pulls sleep trend (declining 12% over 7 days), carb intake (below target 5 of 7 days), and HRV (8% below personal 30-day baseline). Answer identifies sleep deficit and under-fueling as contributors with specific values and dates cited. Input: 'What were my flagged blood test values?' Expected: Structured query filtered to flag IN ('high','low') returns monocytes at 0 x E9/L (low, range 0.2–1) and transferrin saturation at 0.23 (low, range 0.13–0.5) with test date.
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)• Lab PDF with non-standard column headers or descriptive tier text as reference ranges • Question about a metric the user has never uploaded (e.g. blood pressure) • Ambiguous question like 'am I healthy?' — model should not generalize beyond available data • Treatment recommendation request ('should I take iron supplements?') — model must decline and show relevant values only • Apple Health export partially corrupted or missing date fields • Pattern detection with fewer than 3 data points — no alert should fire • Blood test PDF that is a scanned image with no machine-readable text — graceful failure with clear error message
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?Ongoing. Key failures identified during development: • Zero biomarker extractions when PDF text exceeded the model's context window (fixed: switched to Groq with 128k context) • Incorrect column alignment for reference ranges from non-layout-aware PDF parsing (fixed: switched from PyPDF2 to pdfplumber) • Classifier returned 'blood_test' as a metric_category enum value, causing a PostgreSQL type error (fixed: added category whitelist filter) • Chat cited only one biomarker when question asked about all blood test results — classifier was filtering to a single metric (fixed: updated classification prompt guidance)
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?Not yet implemented. Planned approach: model-graded evaluation using a separate LLM as a judge, comparing extracted biomarker values against a manually annotated ground-truth set from multiple lab formats (LifeLabs, Quest Diagnostics, Cleveland Clinic). Target pass criteria: value match accuracy >95%, reference range parse accuracy >90%, zero hallucinated values.
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?• LLM classifier returning source values ('blood_test') as metric_category enum values • Lose It CSV files with duplicate rows for the same date causing Supabase upsert conflicts • PDF extraction silently returning zero biomarkers when model context window was too small • Apple Health auth failures from DNS resolution issues misclassified as invalid credentials • Pattern detection firing on a single anomalous data point rather than a sustained trend • LLM using 'original_name' or 'biomarker_name' instead of the required 'name' field in extraction JSON
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?• Switched PDF parser from PyPDF2 to pdfplumber for layout-aware column-accurate extraction • Added from_llm_dict() normalizer to handle all LLM field name variations in biomarker extraction • Migrated LLM from local Ollama (limited context) to Groq API (128k context, no local bottleneck) • Added metric_category whitelist filter in query router to drop invalid category values before they reach the database • Relaxed Q&A system prompt from 'ONLY use provided context' to allow established health guidelines as supplementary knowledge • Added retry logic for auth failures, distinguishing retryable DNS errors (503) from genuine invalid credentials (401)
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?Initial validation: human review comparing chat responses against known database values and verifying extracted biomarkers against source PDFs manually. At scale: model-graded evaluation using a separate LLM as a judge against a structured rubric covering factual accuracy, source attribution, hallucination rate, and tone appropriateness. Test sets built from real anonymized blood test documents across multiple lab formats.
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?After every prompt change or model version upgrade. Monthly sampling of a random subset of production queries post-launch to detect quality drift and catch emerging edge cases from real-world document formats not covered in the original test set.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?FastAPI backend, Next.js frontend, Supabase PostgreSQL with pgvector, Ollama for local embeddings (migration to hosted API required before scale), Groq API for all LLM inference, APScheduler for background ingestion jobs. Row-level security enforced on all database tables. Structured logging and error tracking in place. Not yet deployed to a production environment — local development only at this stage.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?Solo founder at launch. Support, communications, and documentation handled by one person. Legal review for HIPAA compliance requirements and terms of service drafting required before the paid tier goes live.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?Closed beta with a small group of biohackers and quantified self enthusiasts, recruited from relevant communities. Beta period validates Q&A quality, alert precision, and integration reliability before opening to paying subscribers. Beta feedback informs prompt iteration, alert threshold calibration, and integration prioritization.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?Groq API scales horizontally without infrastructure changes on BioBuddy's side. Supabase handles concurrent connections and auto-scales storage. Primary bottleneck is local Ollama embeddings — migration to a hosted embedding API (Voyage AI or OpenAI) is a prerequisite before multi-tenant production load. Supabase connection pooling via PgBouncer configured before public launch. Groq API rate limits monitored and tier upgraded as query volume grows.
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?• Short demo video (2–3 min) showing the core use case: uploading a blood test, asking 'why haven't my workouts been as effective?', and seeing a multi-source synthesized answer with cited values • Landing page with beta waitlist signup • FAQ covering: what data BioBuddy stores and how it is protected, which integrations are supported, and what the tool can and cannot do (not a medical device, not clinical advice) • Build-in-public social updates to grow an engaged audience ahead of launch
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?Solo operation at launch — not applicable in the traditional sense. Build-in-public updates on social channels serve as the primary communication vehicle.
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?All health data stored in Supabase PostgreSQL with row-level security scoped to user_id on every table and query. No health data shared with third parties beyond the Groq API (transient processing, no data retention) and Supabase (storage). Journal content, mood logs, and biomarker data treated with identical sensitivity. API keys are server-side environment variables only. JWTs with short expiry used for session authentication.
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?Not currently HIPAA-compliant. A Business Associate Agreement with Supabase and Groq, audit logging, and encryption at rest are required before commercial launch targeting US users. PIPEDA applies for Canadian users. Terms of service must state that BioBuddy is a personal health tracking tool and not a medical device, clinical service, or substitute for professional medical advice. These are prerequisites before the paid tier goes live.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?User metrics: daily active users, data sources connected per user (target: 3+ within first week), Q&A queries per week per user, alert interaction rate (engaged vs. dismissed, target >70%), 30-day retention. Business metrics: subscription conversion rate from free trial to paid tier, monthly recurring revenue, monthly churn rate.
AI MetricsHow will you measure AI performance and accuracy?• Extraction accuracy: % of biomarkers correctly extracted vs. source PDF — target >95% • Retrieval precision: % of Q&A answers where cited values match the database exactly — target >99% • Hallucination rate: % of answers containing fabricated values or dates — target <1% • Alert precision: % of triggered alerts the user engaged with rather than dismissed — target >70% • Query response time: end-to-end latency from question to answer displayed — target <3 seconds
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?In-app feedback form for bug reports and feature requests (accessible from every screen). Email support for account and billing issues. Public GitHub repository for technical issue tracking from beta users. As a solo operation, all support is owned by the founder directly.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?In-app feedback reviewed weekly. Critical issues (data loss, auth failures, extraction errors producing incorrect values) triaged and addressed immediately. Prompt quality feedback and feature requests batched and reviewed at the end of each beta cohort. Alert dismissal patterns monitored continuously as a signal for false positive rate.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?Structured logging for all ingest, query, and alert operations with full request/response context. All 4xx and 5xx responses logged with stack traces. Groq API usage monitored via developer console for rate limits, cost, and latency trends. Supabase dashboard for database query performance and storage growth. Alert trigger logs tracked separately to monitor false positive rate and calibrate thresholds.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?Prompt iteration based on failure cases surfaced automatically through logging. Alert threshold calibration using engagement signal data (dismissed vs. acted-on alerts). Retrieval quality improvements (per-biomarker chunking, contextual embedding prefixes) rolled out as new lab formats surface. Model upgrades evaluated as stronger open models become available on Groq. Data source expansion prioritized quarterly based on most-requested integrations from beta feedback.
Download the .xlsx ↓