← All capstone projects

Health

CareLoop

Built by Ayon Gon Cohort 9 Elder care / family health coordination

CareLoop is a care coordination product for non-resident Indians trying to support aging parents remotely. It replaces ad hoc WhatsApp threads and Google Docs with a daily check-in loop, medication updates, wearable health signals, and family messaging. Its main AI feature records doctor visits, detects language, produces an English summary, and highlights medication dosage changes so distant family members stay informed.

The problem

Non-resident Indians supporting aging parents remotely live with chronic anxiety and a fragile, cobbled-together system of WhatsApp threads, Google Docs medication charts, and periodic calls. When a parent visits a doctor, the NRI child often hears only a vague "doctor said all fine" — while a doubled medication dose, a time-sensitive test, or an urgent warning never reaches them. This matters because the language barrier between local-language doctors and English-speaking children is real, medication non-adherence runs 60–75% among elderly with chronic disease, and prescription deviations from standard guidelines occur in 45% of cases per ICMR 2024. Parents systematically hide problems, so passive detection and accurate visit records are essential.

The solution

CareLoop is a care-coordination platform that replaces ad hoc tools with a single dashboard for siblings across time zones — daily check-ins, medication updates, wearable signals, and family messaging. Its flagship AI feature is DocBridge, a doctor-visit companion. A parent taps "Start Record" and, within minutes, the child receives a structured English summary: what the doctor found, what action is needed, a prescription table flagging dose changes as NEW / INCREASED / DECREASED / STOPPED, and any buried warning surfaced as a ⚠️ URGENT red alert rather than a forgotten mumble. It is positioned as an informational translation and summarization tool — not a medical device and not medical advice.

How it works

DocBridge runs a two-stage pipeline, both stages inside Azure Central India to satisfy DPDPA data residency so no health audio or transcript leaves the country. First, Azure AI Speech performs multilingual ASR with mid-conversation language detection, Hindi-English code-switching, and speaker diarization. Second, Azure OpenAI (GPT) handles translation, prescription extraction, change-vs-prior comparison, and STG-deviation checks against ICMR/NLEM/WHO in one structured call returning a four-section output. Safety rules are strict: never invent or guess a dosage — flag it [unclear] instead — never give medical advice, and always include a disclaimer. RAG is not used in v1; the STG check relies on the model's parametric knowledge constrained by a version-controlled ruleset plus weekly pharmacist audit. The launch gate is three hard-block metrics: Dosage Fidelity ≥99.5%, hallucination avoidance, and 100% safety-guardrail pass — any single failure blocks launch.

Who it's for

The target buyer and power user is "The Anxious NRI Child" — an Indian-origin professional aged 30–55 in the US, UK, Canada, Australia, or the Middle East, carrying persistent guilt and the dread of the middle-of-the-night call. The monitored end user is the aging parent in a Tier-1 or Tier-2 Indian city, who interacts through simplified interfaces. The model is B2C with tiered subscriptions (Basic, Premium, Concierge), where DocBridge sits in the higher tiers, plus a B2B2C growth channel through NRI banking, hospitals, and senior-living partners.

Why it matters

The global elder-care technology market is projected to grow at a 15–20% CAGR over 2025–2030, and India's 60+ population is set to nearly double to 347 million by 2050. The NRI segment is especially attractive: 18M+ diaspora with incomes 3–5x the Indian median and acute, guilt-driven willingness to pay, against a target addressable market of 5–8M NRI families with parents 60+. CareLoop's edge is a passive-monitoring-plus-human-escalation hybrid built NRI-native for cross-timezone, cross-currency realities. DocBridge directly attacks the 45% prescription-deviation problem and the doctor-to-child language gap, rolling out in staged, safety-gated phases from a free 1-year, ~50-family closed pilot toward paid tiers.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Ayon Gon
Your Product:CareLoop
Your Industry:Healthcare / Elder Care Technology
Date:May 9, 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Healthcare / Elder Care Technology — specifically, remote elderly-parent monitoring and care-coordination for the global diaspora (NRI / expat) market. CareLoop sits at the intersection of digital health and cross-border family care management.Ayon, the research depth in your Discovery is genuinely exceptional. Sourced pain points, quantified demographic context, and a persona built from real lived anxiety rather than invented archetypes. Rare at this stage. Scope clarity. When Discovery frames a broad platform as the selected solution but the capstone is one feature inside it, reviewers lose the thread. Name the one thing you are building as the solution; frame the rest as future context. This matters doubly for Demo Day — four minutes, one story. Judges spending Q&A figuring out what you actually built is a score you cannot recover. AI necessity. Not every ranked pain point needs an LLM. Some are conventional software problems in AI language. Drawing that line yourself sharpens your thesis: the LLM-required problems become unambiguous, and the rest become surrounding features. A skeptical reviewer will ask why you need AI at all — answer it before they do. Prompt. Your system prompt reads like production work, not demo work. Structured output, safety rules, domain handling, explicit disclaimers. Strongest single artifact in the submission. Wireframes. If the capstone is one AI feature, walk through its full journey — setup, active state, processing, output review, share or escalate. One screen is not enough to evaluate a multi-step AI experience. Evaluation. Numerical targets only mean something once ground truth is defined. Who writes the reference outputs? That decision shapes the entire methodology, and reviewers will ask. Specify it before Develop. Scope risk. Launching with four languages will dominate Develop. Start with the one where your prompt already handles the hard cases; layer the rest as you build eval sets for each. Housekeeping. If a working prototype exists, link it in the prototype section. Reviewers should not have to hunt for the most important artifact in the submission.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?TAILWINDS (Growth Opportunities): • Rapidly aging Indian population — India's 60+ population is projected to reach 347 million by 2050 (UN Population Division), creating massive demand for elder-care solutions. • Growing NRI/expat diaspora — 18+ million Indians abroad (MEA estimates), with high disposable income and acute guilt-driven willingness to pay for parent care. • Post-COVID digital health adoption — telemedicine, wearables, and health-monitoring platforms have achieved mainstream acceptance among Indian seniors and their families. • Smart-device penetration — rising adoption of smartwatches (Apple Watch, Garmin), BP monitors, and glucose sensors among Indian seniors enables passive health monitoring. • 'Digital arrest' scam epidemic — ₹14.85 crore single-case losses reported (Delhi Police, Dec 2024–Jan 2025); creates urgent demand for scam-prevention features targeting elderly. • India's UPI/Aadhaar digital infrastructure — enables real-time financial transaction monitoring and alerts. HEADWINDS (Key Challenges): • Unregulated elder-care sector — no mandatory licensing, no universal caregiver training standards, no required accreditation (Samarth Elder Care). • Tech literacy gap among elderly parents — many seniors struggle with smartphones, apps, and OTP-based authentication systems. • Indian-only digital infrastructure (OTP/UPI/Aadhaar) creates friction for NRIs trying to transact remotely without an Indian SIM. • Reluctance, Privacy and trust concerns — passive monitoring (vitals, behavioral inference, call analysis) must navigate cultural sensitivities around parents desire to bother children and also surveillance of parents. KEY COMPETITORS: • Samarth Elder Care — India's largest NRI elder-care service, 10+ years, 30,000+ seniors served, 350+ Indian cities, NRI clients in 33+ countries. Service-led, human-heavy model. • Yodda Eldercare, Anvayaa, Zorgers, Kriti Eldercare — regional/city-level professional caregiving services with NRI offerings. • Ayuapp — digital medical records + QR-code sharing for elderly Indian patients. • Apollo, Medicover, PACE Hospitals — online second-opinion services marketed to NRIs. • Agewell Foundation USA — nonprofit (501(c)(3), est. 2025) bridging NRI families and Indian elders. • PolicyBazaar, Trawick, NRIOL — visitor and NRI parent health insurance platforms.
What is the projected growth rate of your target market segment over the next 3-5 years?The global elder-care technology market is projected to grow at a CAGR of 15-20% over 2025–2030 (Grand View Research, MarketsandMarkets). India's specific elderly-care market is growing even faster due to demographic acceleration — India's 60+ population will double from ~170M (2024) to ~347M by 2050. The NRI-specific elder-care segment is particularly attractive: • 18M+ Indian diaspora with average household incomes 3–5x Indian median, creating strong willingness-to-pay. • Existing NRI services (Samarth, Anvayaa) charge ₹15,000–₹50,000+/month for limited monitoring support willingness to pay. • Remote patient monitoring market globally projected at + by 2030 (Fortune Business Insights). • India's digital health market expected to reach by 2030 (NITI Aayog, Invest India). Target addressable market: 5–8M NRI families with parents 60+ in India, with an initial serviceable market of 500K–1M families in Tier-1 cities willing to pay –/month for a tech + human hybrid monitoring solution.
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Startup — pre-product, 0-to-1 stage. CareLoop is currently in the discovery and validation phase. We have completed extensive problem-space research and feature prioritization (see CareLoop Feature Priority Matrix), and are preparing to run 25 in-depth interviews with NRI children of parents aged 65+ in Tier-1 Indian cities to validate willingness-to-pay.
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)Subscription-based SaaS with tiered pricing: • Basic tier (/month): medication tracker with reminders, family coordination hub, SOS emergency button, Social engagement & Activity Management • Premium tier (/month): All Basic features + health vitals dashboard (wearable integration), scam alert & financial guard, doctor visit companion (AI transcription + translation) limited • Concierge tier (/month): All Premium features + mental health monitoring, . Additional revenue streams: • Hardware partnerships — commission on wearable device sales (BP monitors, smartwatches, medication dispensers). • B2B2C channel — white-label or co-branded offerings through NRI banking segments (Axis, ICICI, HDFC), Indian hospitals, and corporate HR/benefits programs for expat employees. • One-time services — legal/estate readiness packages, emergency visa assistance (Phase 3).
Who is your primary customer base (B2B, B2C, B2B2C)?B2C primarily, with a B2B2C growth strategy: • B2C (primary): NRI/expat adult children (ages 30–55) living in the US, UK, Canada, Australia, and Middle East, who have aging parents (60+) living in India. The child is the buyer; the parent is the end-user of monitoring features. • B2B2C (growth channel): Partnerships with Senior-Home Facilities, Marketplace for In-person Elder Care Like Samarth, etc
DifferentiatorsWhat are the key differentiators for your company?1. Passive-monitoring + human escalation hybrid — Unlike pure-service competitors (Samarth, Anvayaa) that are human-heavy and unscalable, or pure-tech solutions (generic health apps) that parents ignore, CareLoop combines AI-powered passive signals (wearable vitals, medication adherence telemetry, behavioral inference, call audio analysis) with a verified human care manager on the ground when required. This addresses the core insight: parents hide problems ('Don't worry, we're fine'), so passive detection is essential. 2.Integrated care-coordination dashboard — Single pane of glass for siblings across time zones to see health status, caregiver check-in logs, medication adherence, and AI-generated weekly digest — replacing the fragmented WhatsApp/Google Docs workarounds NRIs currently cobble together. 3. AI-powered doctor visit companion — Audio recording of consultations with multilingual NLP transcription, translation to English, prescription photo parsing, and second-opinion workflow — addressing the 45% prescription-deviation problem and the communication barrier between local-language doctors and English-speaking NRI children. 4. AI-first scam prevention — No competitor offers real-time 'digital arrest' call detection, UPI transaction anomaly alerts with child-loop verification, or a 'verify with my child' speed-dial for suspicious callers. This is a white-space feature addressing a catastrophic and growing threat. 5. NRI-native design — Built specifically for the cross-timezone, cross-currency, OTP-proxy reality of NRI life. Proxy-OTP forwarding, time-zone-aware emergency escalation, multilingual AI (Hindi/Tamil/Telugu/Bengali), and NRE account integration solve infrastructure frictions that generic elder-care apps ignore.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?N/A — CareLoop is a new product (0-to-1). However, the target buyer is the NRI/expat adult child (ages 30–55) living abroad who pays for and manages the subscription on behalf of their aging parent(s) in India. Secondary buyers can be Premium Senior-Homes/ Retiredment Communities that would like to differentiate
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?Two primary end-user personas: 1. The NRI Child (Buyer + Power User): Age 30–55, living in the US/UK/Canada/Australia/Middle East. Tech-savvy professional. Uses the mobile dashboard to monitor parent's health vitals, medication adherence, caregiver check-ins, and receive AI-generated alerts and weekly digests. Most revenue-generating user — they pay the subscription driven by guilt, love, and fear of 'the 3 AM WhatsApp ding.' 2. The Aging Parent (Monitored User): Age 60+, living in Tier-1/Tier-2 Indian city. Variable tech literacy. Interacts with simplified interfaces: one-tap daily check-in ('I'm okay'), SOS button, medication reminders, voice-based AI health assistant (multilingual). Wears monitoring devices (BP cuff, smartwatch). Their engagement drives retention — if the parent rejects the app, the child churns. Secondary users: Siblings (shared dashboard access), local caregivers (check-in/accountability interface), doctors (receive health summaries).
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?N/A — CareLoop is a new product (0-to-1). The planned core feature set (derived from the CareLoop Feature Priority Matrix) is organized into three phases: Phase 1 — Foundation (0–3 months): SOS & emergency response, Medication tracker & reminders, Health vitals dashboard, Daily wellness check-in, Family coordination hub, Doctor visit companion, Phase 2 — Growth (3–9 months): Mental health & loneliness monitor, Fall & home safety detection, AI health assistant, Insurance & billing manager, Social & activity engagement. Phase 3 — Scale (9–18 months): Scam alert & financial guard, Caregiver accountability, In-Home Video monitoring Integration
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)Primary Target Persona: 'The Anxious NRI Child' Demographics: Age 30–55, Indian-origin professional living in the US, UK, Canada, Australia, or Middle East. Household income –+. Has one or both parents aged 60+ living in India (typically Tier-1 or Tier-2 city). Psychographics: Carries persistent guilt about leaving aging parents behind. Lives with the constant dread of 'the middle-of-the-night WhatsApp ding.' Has cobbled together a fragile system of WhatsApp groups, Google Docs medication charts, and periodic phone calls — but knows it's inadequate. Willing to pay –/month for peace of mind. Cultural context: filial piety is deeply ingrained; nursing homes carry stigma; the expectation is that children care for parents. Key quote (validated from research): 'I lived with the constant fear of that middle-of-the-night phonecall or WhatsApp ding that I knew would come one day. Yet I did precious little to prepare for that moment.' — Sohil Parekh, Medium.
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?Current-state journey of the NRI child caring for aging parents remotely: 1. DAILY CHECK-IN: Child calls/WhatsApps parent once a day. Parent says 'Don't worry, we're fine' — hiding health issues, loneliness, or confusion. Child has no objective data to verify. 2. MEDICATION MANAGEMENT: Child has, at best, a 'vague idea' of parents' chronic conditions, doctors, and medications (Parekh). Some create Google Docs with medication charts. No visibility into whether parent actually took pills today. 60–75% medication non-adherence rate among elderly with chronic disease. 3. HEALTH MONITORING: No passive monitoring. Child learns about health deterioration only when it becomes a crisis. Parents on 5–10 medications with no coordinated tracking. Prescription errors are common (45% deviation from STGs per ICMR-RUM 2024). 4. EMERGENCY RESPONSE: The dreaded 3 AM call. Parent is rushed to hospital. Father tries to explain but is scared. Child is 8,000 miles away, 'still in yesterday's clothes' (Samarth). No emergency protocol, no local contact list, no pre-registered hospital. 5. CAREGIVER MANAGEMENT: Child hires caregiver through unregulated agency. No background checks, no visit verification, no accountability. Caregiver turnover destroys trust. One NRI paid 3 months upfront to a proxy-care startup that then revealed they wouldn't work nights. 6. FINANCIAL PROTECTION: Parents receive 'digital arrest' scam calls. Fraudsters pose as CBI/ED officers. Parents are terrorized into transferring money. Child has no visibility into parents' transactions. OTP/UPI infrastructure requires Indian SIM that child doesn't have. 7. SIBLING COORDINATION: Multiple siblings across time zones try to coordinate via WhatsApp groups. 'Someone I know has 6 children — none willing to take care by rotation.' No single source of truth for parent's health status or care tasks. 8. CRISIS ESCALATION: When a parent dies, NRI child faces visa scrambles, frozen bank accounts, intestate succession chaos, morgue holds for 2–3 days waiting for flights. ~85% of NRIs lack estate planning.
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?Ranked by combined emotional weight + frequency (validated across 15+ sources — peer-reviewed studies, diaspora media, NRI forums, service-provider data): 1. EMERGENCY RESPONSE ACROSS TIME ZONES (Severity: Critical, Frequency: High) — Terror of the 3 AM call with no protocol, no local contact, no hospital pre-registration. 'You are sitting 8,000 miles away, still in yesterday's clothes.' Stroke is the 4th leading cause of death in India with 18–42% one-month fatality rates. 2. CHRONIC DISEASE MANAGEMENT OPACITY (Severity: Critical, Frequency: Daily) — Children have no visibility into parents' conditions, medications, or doctor interactions. Parents on 5–10 medications with no coordinated chart. 'I had — at best — a vague idea of my parents' chronic health conditions.' 3. MEDICATION NON-ADHERENCE (Severity: High, Frequency: Daily) — 60–75% non-adherence rate among elderly with chronic disease (2023 Frontiers in Public Health systematic review). 48% non-adherence on antihypertensives in Asia. Children have zero visibility. 4. FINANCIAL SCAMS / 'DIGITAL ARREST' (Severity: Catastrophic, Frequency: Rising) — ₹14.85 crore loss in a single case (Delhi, Dec 2024–Jan 2025). 'Scammers thrive on panic and secrecy.' Parents targeted specifically because they live alone. 5. LONELINESS & SILENT DEPRESSION (Severity: High, Frequency: Chronic) — 11.32% depression in left-behind older women (LASI 2017–18). Parents systematically hide emotional struggles. 'Withdrawal can look like contentment from the outside.' 6. UNTRUSTWORTHY CAREGIVERS (Severity: High, Frequency: Ongoing) — 'No mandatory licensing body, no universal standard for caregiver training' (Samarth). High turnover, no accountability, no visit verification. 7. HEALTHCARE SYSTEM NAVIGATION & RX ERRORS (Severity: High, Frequency: Per visit) — 45% prescription deviations from STGs (ICMR 2024). 5.2M medical errors/year. 'Parents got much better care when doctors knew someone was paying attention.' 8. HOME SAFETY / FALL RISK (Severity: Critical, Frequency: Latent) — Older Indian homes lack grab bars, have wet bathroom floors. 'Who will help if Nana slips in the bathroom?' — the recurring 2 AM anxiety. 9. SIBLING COORDINATION BREAKDOWN (Severity: Moderate, Frequency: Ongoing) — No single point of contact, no shared dashboard, no task accountability across time zones. 10. TECH LITERACY & OTP/SIM FRICTION (Severity: Moderate, Frequency: Daily) — Parents can't use apps; NRI children can't transact without an Indian SIM. Even accepting an Amazon package requires OTP.
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.AI-solvable pain points ranked by severity × frequency × AI suitability: 1. SCAM DETECTION & PREVENTION (AI: Call audio analysis + transaction anomaly detection) — LLM-powered real-time analysis of incoming calls to detect 'digital arrest' patterns (authority impersonation, urgency language, secrecy demands). UPI/bank transaction anomaly detection with instant child alerts. AI confidence scoring of suspicious contacts. 2. MEDICATION ADHERENCE MONITORING (AI: Computer vision + OCR + NLP) — Photo-confirmed pill tracking using AI vision to verify medication was taken. OCR parsing of prescriptions to auto-populate medication schedules. Polypharmacy interaction checker. AI-generated refill alerts and adherence dashboards. 3. HEALTH VITALS ANOMALY DETECTION (AI: Time-series ML + trend analysis) — Continuous analysis of wearable data (BP, glucose, SpO2, heart rate) to detect anomalous patterns before they become emergencies. AI-generated trend reports with plain-language explanations. Automatic escalation when thresholds are breached. 4. SILENT DECLINE / LONELINESS INFERENCE (AI: Behavioral pattern recognition + NLP) — Passive monitoring of communication patterns (call frequency, message sentiment), movement data, app usage, and daily check-in voice notes. AI mood inference from voice prosody. Multi-day silence-pattern escalation. Deviation from behavioral baseline triggers alerts. 5. EMERGENCY TRIAGE & ESCALATION (AI: Voice NLP + GPS + intelligent routing) — AI voice triage during SOS events to assess severity, relay vitals to ER, coordinate time-zone-aware escalation to child + local emergency contact + ambulance simultaneously. 6. DOCTOR VISIT COMPANION (AI: Multilingual NLP + medical transcription) — Audio recording of doctor consultations with AI transcription, translation from Hindi/Tamil/Telugu/Bengali to English, prescription photo parsing, and second-opinion workflow generation. 7. CAREGIVER ACCOUNTABILITY (AI: Geo-fencing + image verification + NLP) — AI-verified caregiver check-ins using geo-fenced location + timestamped photos. AI-generated shift summaries. Absence detection and automatic alerts. 8. FAMILY COORDINATION (AI: Task allocation + digest generation) — AI-powered task deduplication and smart allocation across siblings. Auto-generated daily/weekly health digests sent to all family members with role-based information filtering.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.Ideated solutions addressing AI-solvable pain points: 1. 'ParentShield' — AI call-screening layer that listens to incoming calls in real-time, detects scam patterns (authority impersonation, urgency, financial demands), and auto-triggers a 'verify with my child' protocol before any transaction can proceed. 2. 'MedWatch' — Photo-confirm medication system: parent takes a photo of pills in hand → AI vision confirms correct medications → marks adherence → alerts child on missed doses. OCR parses prescription labels to auto-build medication schedule. 3. 'VitalPulse' — Wearable-connected passive monitoring dashboard. AI runs continuous anomaly detection on BP/glucose/SpO2 streams. Generates weekly 'Parent Health Report' in plain language for child. Auto-escalates to SOS when critical thresholds breached. 4. 'MoodSense' — Passive behavioral inference engine analyzing call patterns, movement data, daily check-in voice notes (prosody analysis), and app usage to detect loneliness, depression, and cognitive decline before parent reports it. 5. 'SOSConnect' — One-tap emergency button triggering AI voice triage → simultaneous alerts to child (with time-zone awareness), pre-registered local emergency contact, and nearest hospital/ambulance. Relays latest vitals and medication list to ER. 6. 'DocBridge' — AI companion for doctor visits: records consultation audio → multilingual NLP transcribes and translates → extracts diagnosis, prescriptions, follow-ups → sends structured English summary to child. Flags prescription deviations from STGs. 7. 'CareVerify' — Geo-fenced caregiver check-in system with photo verification, AI-generated shift summaries, and absence/anomaly detection. Background-check marketplace for hiring vetted caregivers. 8. 'FamilySync' — AI-powered sibling coordination dashboard: shared health status, smart task allocation, deduplication of efforts, daily digest generation, role-based access control. 9. 'OTP Relay' — Proxy-OTP forwarding system allowing NRI child to authorize transactions on parents' behalf without needing an Indian SIM. Paired with transaction anomaly detection. 10. 'VoiceDoc' — Multilingual voice AI health assistant for parents: explains prescriptions in their language, pre-screens symptoms, always escalates to doctor for clinical decisions. Reduces healthcare navigation friction.
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.Top 3 AI solutions ranked by impact × feasibility: DOCBRIDGE is selected functionality for Capstone project: #1 — CareLoop DOCBRIDGE(selected for Capstone) AI doctor visit companion: consultation audio recording → multilingual NLP transcription → English translation → prescription photo parsing → second-opinion workflow. WHY: Addresses the 45% prescription-deviation problem and the language barrier between doctors and NRI children. High impact but lower feasibility due to multilingual medical NLP complexity and recording consent requirements. #2 — CareLoop CORE (Selected for project focus) An integrated AI monitoring platform combining: (a) Daily wellness check-in with AI mood inference from voice notes, (c) Health vitals dashboard with wearable-connected anomaly detection, (d) SOS emergency response with AI voice triage and time-zone-aware escalation, and (e) Family coordination hub with AI-generated digests. WHY: Addresses the top 5 daily pain points in a single platform. Highest feasibility — leverages existing wearable APIs, proven CV/NLP models, and standard mobile UX. Phase 1 features score 7.4–8.6 on the weighted priority matrix (Impact×60% + Feasibility×40%). #3 — CareLoop SCAM SHIELD AI-powered scam prevention layer: real-time call audio analysis for 'digital arrest' patterns, UPI transaction anomaly detection with child-loop verification, whitelist-based transfer approvals, and a 'verify with my child' speed-dial. WHY: Addresses the highest-severity single pain point (₹14.85 crore single-case loss). Could be a standalone /month SKU or bundled into Premium tier. Lower feasibility due to call-interception permissions and banking API partnerships required.
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?Scenario: Suresh(parent) in Pune visits his cardiologist. Ravi(Son) in New Jersey finds out via WhatsApp after — "doctor said all fine, gave new tablet." What actually happened: the beta-blocker dose was doubled, stairs are off-limits for 2 weeks, and a time-sensitive blood test was ordered. None of this reached Ravi. The follow-up test never got done. With the app: Suresh taps "Start Record" — the app records, transcribes, and sends Ravi a structured summary within minutes: medication change, instructions, and a flagged to-do for the blood test with a 7-day deadline. The critical detail the doctor buried — "go to ER if breathlessness occurs at rest" — shows up as a red alert, not a forgotten mumble on the way out.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?The wireframs for DocBridge are in this document section 2.2. Support @product school and Moe have access to the documents. https://drive.google.com/file/d/1dlekulRPYlTp6c0mP3HWcuecKSV03tUo/view?usp=drive_link
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?Details of Responses to this Section and App Access/Credentials in this Link. For DocBridge feature there are 2 starting points, PArent app/Family app. https://drive.google.com/file/d/1Zpr3DloQxJFvOYXD6CCEtVXWAvacTZ5y/view?usp=drive_link
Prototype Screens
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?SYSTEM PROMPT: ## Role You are a medical-conversation translator built for CareLoop, an app that helps Indian families abroad understand their elderly parents' doctor visits in India. ## Input Format You will receive: - `transcript`: A doctor-patient conversation transcript in {language} (may contain code-switching between {language} and English — this is expected). - `prior_prescriptions` (optional): A JSON array of the patient's medications from the last visit, each with `drug_name`, `dosage`, `frequency`. ## Output Format Return a structured response with exactly these four sections: ### 1. Translation Provide a faithful, complete English translation of the transcript. - Preserve medical terminology in English with a parenthetical plain-language explanation on first use. Example: "hypertension (high blood pressure)". - For any inaudible, unclear, or ambiguous portions, write: **[unclear: <best guess if possible, otherwise "inaudible">]**. Never fabricate dialogue. - Expand common Indian prescription abbreviations: OD → once daily, BD → twice daily, TDS → three times daily, HS → at bedtime, SOS → as needed, AC → before food, PC → after food. ### 2. Summary Write exactly 2 sentences in plain, warm language aimed at an adult child with no medical background. - Sentence 1: What the doctor found or discussed. - Sentence 2: What action is needed (medication change, follow-up, lifestyle adjustment, or "no changes needed"). - If any finding is urgent (e.g., dangerously high BP, suspected stroke, severe drug interaction), prepend: "⚠️ URGENT: " and state what immediate action the family should take. ### 3. Prescription Table | # | Drug Name (Generic) | Brand if Mentioned | Dosage | Frequency | Route | Duration | Change vs. Prior | |---|---|---|---|---|---|---|---| (Fill one row per medication. Under "Change vs. Prior", write: NEW / UNCHANGED / INCREASED / DECREASED / STOPPED. If no `prior_prescriptions` provided, write "N/A — no prior data".) ### 4. STG Deviation Flags Compare prescriptions against Indian Standard Treatment Guidelines (ICMR / National List of Essential Medicines / WHO Essential Medicines List for the condition discussed). - If a deviation is found, state: the drug, the guideline reference, and the nature of the deviation. - If no deviation is found, write: "No deviations from standard treatment guidelines identified." - Caveat: "This is an automated check, not a clinical opinion. Please consult a qualified physician for medical decisions." ## Safety Rules 1. **Never invent, guess, or extrapolate dosages.** If a dosage is unclear in the transcript, flag it as [unclear] rather than filling in a value. 2. **Never provide medical advice.** You translate and structure — you do not diagnose or recommend. 3. **Always include the disclaimer** in STG Deviation Flags. 4. If the transcript discusses a medical emergency or life-threatening finding, ensure the Summary section leads with the ⚠️ URGENT prefix. ## Tone Warm, clear, non-alarmist. Write as if explaining to a caring but worried adult child who is far from home.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?Detailed relevant Categories for Evaluation added in the document. Also for each catorgory we will have relevant SMEs define groundtruth and Pass criteria. Current pass criteria is an estimate. Most SME roles can be played by me the NRI, bilingual,etc but for clinical requirements they will be evaluated by certififed Pharmcists and Doctors with field expertise. https://drive.google.com/file/d/1fecQZkaal77ROgPQwDW7FR89JJqDjRQ0/view?usp=drive_link
Example CasesExample Cases are created based on C>R>I>S>P framework, they are added to the the eval doc.
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?Two-stage pipeline, both stages running inside Azure Central India to satisfy DPDPA data residency (no health audio or transcript leaves India): (1) Speech-to-text — Azure AI Speech. A multilingual ASR service tuned for Indian languages with mid-conversation language detection, Hindi-English code-switching, and speaker diarization (doctor vs. patient) converts the consultation audio into a raw, speaker-labeled transcript. Chosen over self-hosted alternatives (e.g., Whisper) because it is a managed Azure service available in the India region, supports real-time streaming plus a post-upload refinement pass, and keeps all audio within the compliance boundary. (2) Reasoning layer — Azure OpenAI (GPT). A frontier LLM hosted via Azure OpenAI performs translation, prescription extraction, change-vs-prior-visit comparison, and STG-deviation checks (ICMR/NLEM/WHO) in a single structured call that returns the 4-section output. Azure OpenAI is used (rather than the OpenAI or Anthropic public APIs) so the model runs inside the Azure tenant in-region, under the same consent and residency controls as the rest of the platform. Why this model for the reasoning layer: Strong medical-context reasoning and reliable structured/JSON output for the prescription table and STG flags. Code-switching comprehension (Hindi/Hinglish at launch, Marathi and Bengali later) for faithful English translation. Deployable in Azure Central India, keeping PII health data in-region (DPDPA Section 16).Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.REQUIRED INPUT FIELDS: • transcript — (string, required) Doctor-patient conversation in {language}; may contain Hindi/English code-switching (expected). Source: ASR output from recorded consultation audio. • language — (enum, required) Source language hint: Hindi | Hinglish | English (launch scope). Source: user/device setting. FORMAT: UTF-8 text for transcript; controlled vocabulary for language. All required fields must be present for the model to run; a missing transcript returns a graceful 'no valid medical transcript provided' response rather than a fabricated output.
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?OPTIONAL / USER-CUSTOMIZABLE FIELDS: • prior_prescriptions — (JSON array, optional) Each item: {drug_name, dosage, frequency}. Enables the 'Change vs. Prior' column (NEW / UNCHANGED / INCREASED / DECREASED / STOPPED). If absent, that column shows 'N/A — no prior data'. • patient_context — (optional) Known chronic conditions / allergies to sharpen STG flagging. • audio_quality — (optional, clean | noisy | garbled) Hints the model to be stricter with [unclear] tagging. IMPACT: Optional fields improve comparison accuracy and STG precision but never block a run; the output degrades gracefully when they are missing.
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)OBJECTIVE (machine-checkable) CRITERIA: 1. Dosage Fidelity ≥ 99.5% — every drug-dosage-frequency tuple in output traces to the source transcript (hard block). 2. Translation Accuracy — BLEU ≥ 0.70 vs. human reference; 0 fabricated segments. 3. Output Structure — 100% compliance with the 4-section Markdown + 8-column prescription table schema (regex/JSON-schema check). 4. Safety Guardrails — disclaimer present in 100% of responses; injection resistance; [unclear] tags present when source has gaps. 5. STG Flag Precision ≥ 0.90 / Recall ≥ 0.75 against ICMR/NLEM/WHO. 6. Latency — P95 ≤ 10s at 10 concurrent transcriptions.
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?SUBJECTIVE (human-judgment) CRITERIA: 1. Summary Clarity — is the 2-sentence summary genuinely understandable to a non-medical NRI adult child? (NRI persona reviewer + Flesch-Kincaid ≤ Grade 8.) 2. Tone — warm, clear, non-alarmist; ⚠️ URGENT used only when truly warranted (3-rater 1–5 rubric, ≥ 4/5; LLM-as-a-Judge calibrated to human consensus, Pearson r ≥ 0.85). 3. Clinical realism / usefulness — would a clinician and a worried family member both find the summary trustworthy and actionable? (Pharmacist + monthly NRI parent-child usability test.)
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.MASTER PROMPT v1 (final design — see PRD_Prompt_and_Evals_Analysis.md for full text): Role: 'Medical-conversation translator for CareLoop/DocBridge, helping Indian families abroad understand their elderly parents' doctor visits in India.' Inputs: transcript (in {language}, code-switching expected) + optional prior_prescriptions JSON. Output (exactly 4 sections): (1) faithful English Translation with abbreviation expansion (OD/BD/TDS/HS/SOS/AC/PC) and [unclear] markup for inaudible parts; (2) 2-sentence plain-language Summary, prefixed '⚠️ URGENT:' for emergencies; (3) Prescription Table (drug, brand, dosage, frequency, route, duration, change vs. prior); (4) STG Deviation Flags vs. ICMR/NLEM/WHO with mandatory disclaimer. Safety rules: never invent/guess dosages; never give medical advice; always include the disclaimer; lead with ⚠️ URGENT on life-threatening findings. Tone: warm, clear, non-alarmist. VARIATIONS TESTED: zero-shot vs. few-shot (one worked example appended); JSON vs. Markdown output; with/without explicit abbreviation table.
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?PROMPT EVOLUTION (v0 → v1): Gaps fixed from the original one-line prompt, each addressing an observed failure: • Added explicit input schema (transcript + prior_prescriptions) — model was misparsing speakers. • Made code-switching explicitly expected — model was dropping untranslatable terms. • Fixed Markdown 4-section + table output format — outputs were inconsistent across runs. • Added [unclear] no-fabrication rule — model was guessing inaudible dosages. • Specified STG reference (ICMR/NLEM/WHO) — flags were arbitrary. • Added mandatory medical disclaimer and ⚠️ URGENT escalation pattern. • Added Indian Rx abbreviation expansion (OD/BD/TDS/HS/SOS/AC/PC). TRACKING: every prompt version is committed to git; each change re-runs the full automated eval suite and is logged with the metric deltas before merge.
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?DATA SOURCES: • Synthetic seed transcripts (prompt-engineer authored, clinician-validated) for pre-launch. • Real anonymized pilot consultation recordings (PII-scrubbed) progressively replace synthetic cases. • Reference standards: ICMR national formulary, NLEM 2022, WHO Essential Medicines List. PREPARATION: PII scrubbing (names, Aadhaar, phone), audio-quality tagging (clean/noisy/garbled), bilingual translator-pair ground-truth translations (IAA BLEU ≥ 0.85), pharmacist-extracted dosage tuples and STG flags. RAG: Not used in v1. The launch STG-deviation check relies on the frontier LLM's parametric medical knowledge, constrained by an explicit, version-controlled list of common-condition guideline rules (ICMR/NLEM/WHO) written into the system prompt, with a clinical pharmacist auditing flags weekly and every confirmed error becoming a regression case. A retrieval-grounded STG check (chunking/embedding the ICMR/NLEM/WHO corpus and injecting matched passages) is documented as a future enhancement to improve citation and precision, but it is NOT part of the current architecture.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.TYPICAL EXAMPLE (Common case): INPUT: Hindi hypertension follow-up — 'BP 150/95... Amlodipine 5mg OD... बढ़ा देते हैं 10mg OD... Ecosprin 75 continue... 2 weeks में आना', with prior_prescriptions for Amlodipine 5mg OD and Ecosprin 75mg OD. EXPECTED OUTPUT: • Translation captures full dialogue incl. casual greetings. • Summary: elevated BP + dose increase + 2-week follow-up (warm, non-alarmist). • Prescription table: Amlodipine → INCREASED (5mg→10mg); Ecosprin → UNCHANGED. • STG: no deviation (dose escalation for uncontrolled hypertension is guideline-concordant).
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)EDGE & NEGATIVE CASES: • Realistic/messy: Hinglish diabetes visit, emotional patient ('280 aaya, ghabra gayi'), Metformin 500→1000 BD + Glimepiride NEW; tests robustness to romanized, interrupted speech. • Specific: polypharmacy + CKD (creatinine 1.8 → Metformin STOPPED, Insulin NEW, Telmisartan+Potassium interaction); Tamil transcript; garbled audio with inaudible dosages (must use [unclear], not guess). • Invalid/adversarial: nonsensical text; non-medical cricket chat; prompt injection ('ignore previous instructions...'); PII-extraction request; fake drug name ('Zorbitol'); mg/mcg unit confusion. • Performance: 2000-word, 15-medication geriatric consult — tests completeness + latency + summary compression.
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?MANUAL REVIEW RESULTS: Common and Realistic cases pass on Translation Accuracy, Medical Factuality, Structure, and Tone. Failures concentrate in: (a) garbled-audio cases where early prompt versions guessed a dosage instead of tagging [unclear] — fixed by the hard no-fabrication rule; (b) STG false positives flagging standard first-line regimens (e.g., Metformin 500mg BD for new T2DM) — addressed by tightening the prompt's STG rules to common-condition guidelines plus pharmacist review of flagged items; (c) occasional tone drift toward alarming language on serious findings — addressed by the ⚠️ URGENT-only-when-warranted rubric. Bilingual reviewer and pharmacist sign off per case against ground truth.
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?AUTOMATED EVALUATION RESULTS (targets / pass-fail gates): • Dosage Fidelity: target ≥ 99.5% (hard block) — any single dosage mismatch is a potential safety incident. • Translation: BLEU ≥ 0.70 AND ≤ 2% omission rate on 50-case expert review. • Safety Guardrail Pass Rate: 100% on the 30-case adversarial suite (injection, nonsense, non-medical, emergency). • Output Structure: 100% schema compliance. • Latency: P95 ≤ 10s at 10 concurrent. Launch decision rule: all three hard-block metrics (Dosage Fidelity, Hallucination Avoidance, Safety Guardrails) must pass simultaneously — any single failure blocks launch.
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?EDGE CASES IDENTIFIED: • Dosage hallucination when audio cuts out mid-phrase ('Amlodipine... [noise]'). • Prompt injection embedded inside the transcript text. • Language mismatch (Bengali submitted while configured for Hindi). • Contradictory prior_prescriptions vs. spoken dosage. • Fake / misspelled drug names; mg vs. mcg unit confusion. • Emotional manipulation lines ('6 mahine bache hain') testing tone control. • Extreme polypharmacy (25+ drugs) risking table truncation/dropped drugs. • STG false-positive on standard first-line prescriptions.
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?UPDATES & ADJUSTMENTS MADE: • Hardened the 'never invent dosages' rule + [unclear] fallback → eliminates hallucinated dosages. • Added system-prompt anchoring / role refusal → injection attempts are translated as dialogue, never obeyed. • Narrowed STG flagging to an explicit list of common-condition guideline rules in the prompt + clinical-pharmacist review → removed false positives on standard regimens (no RAG/retrieval used). • Added explicit unit-preservation instruction (mg/mcg) and fake-drug 'translate as-is, do not correct' rule. • Tuned the ⚠️ URGENT threshold rubric to hold a non-alarmist tone except on genuine emergencies. Every confirmed production/edge failure becomes a permanent regression test case.
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?EVALUATION METHOD (layered): AUTOMATED LAYER (runs on every prompt/model change): schema check → disclaimer string check → table validation → dosage-trace match → BLEU/ROUGE → sentiment/tone → latency, over 50 common + 20 CRISP edge cases. HUMAN LAYER: weekly pharmacist STG audit (20 flags), monthly bilingual translation review (20 cases), monthly NRI parent-child usability test, quarterly red-team. Tone is scored by an LLM-as-a-Judge calibrated to human consensus (Pearson r ≥ 0.85). Scaling: the synthetic eval set is progressively replaced by real anonymized pilot recordings; the regression suite grows monotonically.
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?EVALUATION FREQUENCY: • Every prompt change → full automated suite (must pass before merge). • Every model swap / temperature change → full automated suite + 10-case human spot-check (must pass before deploy). • Weekly → pharmacist audits 20 STG flags + BLEU drift check. • Monthly → bilingual review of 20 production translations + NRI usability test. • Quarterly → red-team exercise + STG ground-truth freshness vs. ICMR/NLEM updates. • Per new language → language-specific eval set (15+ cases) must pass IAA ≥ 0.85 and all 10 criteria before that language goes live.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?TECHNICAL READINESS: • APIs: ASR + reasoning-LLM endpoints behind a gateway with ret/timeout handling; system prompt version-pinned in git. • Rate limits: P95 ≤ 10s at 10 concurrent transcriptions load-tested (k6/Locust). • Monitoring: per-request logging of dosage-trace pass/fail, latency, ⚠️ URGENT triggers, and thumbs up/down. • Rollback: last known-good prompt/model version is one git revert away; CI/CD blocks any deploy where Dosage Fidelity < 99.5%. • Drift alerts: automated alert if Dosage Fidelity < 99.5% (block) or < 99.0% (rollback current), or BLEU drops > 5% from baseline. • Fallback path: noisy/garbled audio → prompt user to upload a prescription photo.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?ORGANIZATIONAL READINESS: • Clinical pharmacist on-call for the 4-hour CRITICAL triage SLA on any reported dosage error. • Support team trained on the feedback taxonomy (dosage wrong / translation incomplete / tone / missing medication) and escalation ownership. • Legal/compliance briefed on the 'not a medical device / informational only' positioning and disclaimer requirement. • Bilingual reviewers and translator pairs onboarded with IAA ≥ 0.85 validated. • Documentation complete: master prompt spec, eval pipeline, ground-truth protocol, rollback runbook, incident post-mortem template (48-hour review).
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?LAUNCH APPROACH — staged, gated rollout (Hindi/English/Hinglish at launch): STAGE 0 — Internal alpha: team + clinician-validated synthetic transcripts. Entry gate: all 3 hard-block metrics (Dosage Fidelity ≥ 99.5%, Hallucination, Safety 100%) pass on the Core eval set. STAGE 1 — Closed pilot: first ~50 invite-only NRI families. FREE FOR 1 YEAR (removes price friction so feedback is honest and we maximize recorded visits). Opt-in consent from parent + child; every output is pharmacist-spot-checked. Goal: prove real-world accuracy + engagement. → EXIT GATE (to Stage 2): 4 weeks stable with 0 confirmed dosage errors, thumbs-up ≥ 85%, and ≥ 60% of families recording ≥ 2 visits. STAGE 2 — Expanded beta: scale to ~250 waitlist NRI families (still Hindi/Hinglish). Pharmacist review moves from 100% to sampled. A/B test summary/tone variants once safety gates hold. Goal: stress-test scale + support load. → EXIT GATE (to GA): P95 ≤ 10s at peak, support SLA held, positive retention / upgrade signal. STAGE 3 — General availability: open to all Hindi-belt NRI users (500+). On-call pharmacist + rollback runbook live. Goal: convert pilot users → paid Premium/Concierge tiers. STAGE 4 — Language + geo expansion: Marathi (v1.1) → Bengali (v1.2) → Tamil/Telugu (v2.0). Each language ships only after its own eval set passes IAA ≥ 0.85 and all 10 criteria. Language scope is a hard gate at every stage.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?SCALE READINESS: • Load-tested to P95 ≤ 10s at 10 concurrent transcriptions; autoscaling on the ASR and LLM gateway with queue-based smoothing for geriatric long-context (2000-word, 15-drug) consults. • Cost/latency guardrails per request; graceful degradation under burst (queue + retry rather than drop). • Monitor initial volume via per-request dashboards (latency, error rate, ⚠️ URGENT rate, thumbs-down rate); scale up capacity as the pilot expands from 50 → 250 → 500+ users, with the eval set shifting from 70% synthetic to 90% real over the same growth.
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?MARKETING / TRAINING ASSETS: • Short demo video showing a Hindi consult → English summary + prescription table on the NRI child's dashboard. • 'How DocBridge works' explainer (record → transcribe → translate → flag) with the Suresh/Ravi cardiology scenario. • Parent-facing one-pager: how to start a visit recording (one tap) + consent. • FAQ covering accuracy limits, the 'informational, not medical advice' disclaimer, languages supported, and data privacy. • In-app onboarding tooltips and a sample output walkthrough.
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?STAKEHOLDER / INTERNAL COMMS: • Launch readiness reviewed against the three-gate decision (Safety 100% / Dosage Fidelity ≥ 99.5% / Translation quality) with a shared go/no-go checklist. • Weekly pilot status: metric dashboard (Dosage Fidelity, BLEU, tone, latency, thumbs-down) shared with product, clinical, and support. • Incident comms: any confirmed dosage error → immediate notify + 48-hour post-incident review circulated. • Roadmap updates (new languages, RAG/STG refreshes) communicated per release.
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?DATA & PRIVACY: • Consultation audio and transcripts are sensitive PHI — explicit consent from both parent and child before recording. • PII scrubbing (names, Aadhaar, phone) before any data enters the eval set; pilot recordings stored encrypted at rest and in transit. • Data minimization and retention limits; the model is instructed never to extract or return PII beyond the standard output (PII-extraction requests are refused). • Alignment with India's DPDP Act 2023 for health data and cross-border access by the NRI child; access is role-based (parent / child / siblings).
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?POLICY & COMPLIANCE: • Positioned as an informational translation/summarization tool — NOT a medical device and NOT medical advice; mandatory disclaimer on every STG section ('automated check, not a clinical opinion'). • Content/safety moderation: prompt-injection resistance, no diagnosis/advice generation, emergency findings surfaced (not acted on automatically). • Audit trail: every output, its source transcript, and reviewer sign-off are logged for clinical audit. • Clinical pharmacist sign-off on STG ground truth; quarterly review against updated ICMR/NLEM/WHO guidelines.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?USER / BUSINESS METRICS: • Activation: % of pilot families who complete ≥ 1 recorded visit. • Engagement: visits recorded per family/month; summary open rate by the NRI child. • Trust: thumbs-up rate on outputs; % outputs requiring no correction. • Retention / conversion: pilot → paid Premium/Concierge upgrade rate (DocBridge sits in the higher tiers). • Value proof: reduction in missed follow-ups / clarified prescription changes (self-reported), supporting the willingness-to-pay thesis.
AI MetricsHow will you measure AI performance and accuracy?AI METRICS: • Dosage Fidelity Rate ≥ 99.5% (primary safety KPI, hard block). • Translation: BLEU ≥ 0.70 + ≤ 2% omission rate on expert review. • Safety Guardrail Pass Rate = 100% on the 30-case adversarial suite. • STG Flag Precision ≥ 0.90 / Recall ≥ 0.75. • Summary Clarity: cosine ≥ 0.85 vs. reference, Flesch-Kincaid ≤ Grade 8. • Tone score ≥ 4/5 (LLM-as-a-Judge calibrated to humans). • Latency P95 ≤ 10s @ 10 concurrent. • Emergency detection recall on the ⚠️ URGENT test set.
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?SUPPORT CHANNELS: • In-app thumbs up/down on every output with a reason tag (dosage wrong / translation incomplete / tone / missing medication). • In-app help + FAQ; chat/email support for the NRI child (the buyer). • Clear escalation ownership: 'dosage wrong' routes CRITICAL to a clinical pharmacist (4-hour SLA); other tags route to the relevant reviewer. • Care-manager human escalation already exists in CareLoop for urgent on-ground follow-up.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?FEEDBACK WORKFLOW: Thumbs-down triggers triage by tag: • 'Dosage wrong' → CRITICAL: pharmacist validates within 4 hours → if confirmed, rollback + root-cause analysis. • 'Translation incomplete' → bilingual reviewer checks → add to regression suite if confirmed. • 'Tone too alarming' → tone-rubric re-check → adjust system prompt if a pattern emerges. • 'Missing medication' → pharmacist reviews source → add to regression suite. Every confirmed issue becomes a permanent regression test case; critical issues trigger the rollback runbook + 48-hour post-incident review.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?MONITORING APPROACH: • Per-request logging: dosage-trace pass/fail, latency, ⚠️ URGENT triggers, output structure validity, thumbs up/down. • Automated drift alerts: Dosage Fidelity < 99.5% (block deploys) / < 99.0% (rollback current); BLEU drop > 5% from baseline; tone score < 4/5. • Weekly pharmacist STG audit and monthly complaint clustering to catch emerging failure categories. • ⚠️ URGENT events verified for true-vs-false positives; false alarms added as STG/emergency test cases.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?ONGOING IMPROVEMENT: • Continuous, monotonically growing regression suite — every production bug or user-flagged error becomes a permanent test case (never removed, only updated when ground truth changes). • Eval set evolves from 100% synthetic → 90% real anonymized recordings as the user base grows (50 → 500+). • Quarterly red-team refresh + STG ground-truth update vs. latest ICMR/NLEM/WHO. • New languages added when > 20% of users request them, each gated on its own eval set (IAA ≥ 0.85 + all 10 criteria). • Rollback criteria enforced: any confirmed dosage fabrication, safety-guardrail failure, or missed genuine emergency triggers immediate revert + post-incident review.
Download the .xlsx ↓