← All capstone projects

Food

MyLens

Built by Olivia Lottmann Cohort 9 Consumer meal planning / food tech

MyLens is a meal-planning product for households with real dietary constraints, changing cravings, and a growing personal recipe library. Instead of only filtering meals, it tries to learn a household over time so it can assemble weekly plans from the recipes people already trust. The system pairs AI reasoning for messy preference tradeoffs with deterministic code for constraint checking, grocery-list generation, and plan completeness.

The problem

For couples who love cooking together but resent the planning, the weekly loop breaks down at every stage. Sunday planning gets avoided — and skipped more often than done — which cascades into unoptimized shopping, improvising all week, and food waste. Plans then break mid-week when schedules shift or cravings change. The deeper problem is that existing tools store recipes but never learn a household. 61% of households say meal planning matters, but only 37% actually plan. Ingredient sharing across a week is combinatorially hard for a human, the best version of a dish lives in no one's memory, and point solutions each win only one job — none combine constraint-safe trust, cross-recipe optimization, and adaptive replanning.

The solution

Madeleine is a meal planner built for constraint-heavy couples that learns a household over time rather than just filtering meals. One tap on "What do we eat this week?" produces a full week from the recipes people already trust, respecting dietary constraints, optimizing shared ingredients, and planning around cooking sessions rather than individual meals. Its core differentiator is compounding household taste memory — constraints, signature dishes, validated adaptations, and dislikes that surface over time and become the switching cost. Free-text mood steering ("Asian week," "comfort food") reshapes the plan, and photo-log replanning lets a couple snap what they actually ate so the rest of the week readjusts.

How it works

The guiding principle of the current design is AI for judgment, code for truth. Madeleine runs on Claude Sonnet 4.5 through a Supabase Edge Function, using its vision capability for photo interpretation and recipe import, with a household profile injected as JSON on every call. The master prompt (V1.9) applies library-first rules — scan the saved recipe library before inventing anything — plus hard constraint rules, soft priorities, and a strict JSON schema with structured day/meal "covers" and a recipe-ID contract. Deterministic code then handles everything that must be correct: resolving recipe identity by ID with a fuzzy-match backstop, copying exact ingredients, building the grocery list as the provable sum of the week, and rejecting any plan that would save with missing ingredients. RAG is deliberately deferred until the adaptation memory grows in V2.

Who it's for

Madeleine is a B2C product priced per household, not per individual — the paying unit is the couple. The target segment is dual-working Gen Z / Millennial / Gen X couples who cook and eat together but resent the coordination overhead, where at least one partner has a meaningful dietary constraint like gluten-free, dairy-free, or vegetarian. They batch-cook, shop weekly across stores, get inspired by restaurants and social media but lose the ideas, and have abandoned planning systems before. They aren't looking for a recipe app — they can already cook — but for an extension of their own memory and taste. The reference household is Olivia and Eric in Grenoble, the primary discovery subjects and first testers.

Why it matters

Home cooking is dominant and constraint-aware planning is mainstream: 91% of consumers use online recipes and 43% specifically seek diet-specific options, yet no product owns the full weekly workflow. Charging per household roughly doubles willingness-to-pay per acquisition, and the value compounds — week 12 plans feel more "theirs" than week 1, making the taste profile the moat competitors can't replicate. The business is pre-revenue 0-to-1, validating a single hypothesis: that a constraint-heavy couple given one-tap weekly plans will still be using it after four cycles. Evaluation revealed the central finding — a 91.7% automated pass rate but only 64% human delight — which drove the pivot to splitting creative judgment (the model) from provable correctness (code).

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) Version 1.0
Your Name:Olivia Lottmann
Your Product:Madeleine
Your Industry:Consumer Food Tech
Date:June 9th, 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Consumer Food Tech, in particular AI-assisted household meal planning. Madeleine operates at the intersection of recipe apps, grocery coordination tools and personal nutrition planning. The category is currently served by point solutions (Mealime, Plan to Eat, AnyList, Paprika) that each address one node of a weekly workflow, but no product combines household coordination, constraint-safe planning, taste learning and adaptive mid-week replanning into one friendly system.Please leave this area blank. This space is for the Instructor to provide you with feedback.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?Tailwinds (Opportunities): • 91% of consumers use online recipes and 43% specifically seek diet-specific or healthy options (Chicory 2024 Consumer Survey, chicory.co). • Nearly 9 in 10 US adults eat home-cooked meals multiple times per week, daily home-cooks rate their diets significantly healthier (Pew 2025 Healthy Eating Report). • 69% of Americans say rising healthy-food costs make eating well harder (Pew 2025), which pushes households back to planning home-cooked meals. • Growing cultural shift: Gen Z/Millennial couples share meal planning rather than assigning it to one partner, creating demand for collaborative tools. Headwinds (Challenges): • Category is crowded with free point solutions, so the switching cost for any single feature is low. • Retention is structurally weak: 61% of households say meal planning matters, but only 37% actually plan before cooking (Ncube et al. 2024 household study). The gap between intention and action is the core product challenge, but also a real opportunity to close that gap with Madeleine. • Yummly's shutdown (December 2024) created user distrust, people treat meal tools as durable household infrastructure and fear losing their recipe memory. Competitors and gap analysis: • Mealime: Strongest on guided weeknight planning and dietary filtering. Weakness: closed recipe ecosystem, weak on household learning, adaptation memory, and mid-week replanning. • Plan to Eat: Strongest on planner-centric workflows including leftovers and frozen meals. Weakness: requires heavy up-front curation, more "powerful planning system" than adaptive copilot. • AnyList: Strongest on shared grocery coordination and real-time sync. Weakness: logistics-first, not health/dietary intelligence-first; no adaptive planning or taste learning. • Paprika: Strongest on recipe archive and cross-device organization. Weakness: weak planner-to-shopping connection, no constraint intelligence or dynamic replanning. • Cozi: Strongest on family calendar integration. Weakness: shallow on dietary constraints, nutrition logic, and culinary intelligence. Critical gap: No existing product combines shared household coordination and constraint-safe recipe trust and cross-recipe ingredient optimization and adaptive mid-week replanning and compounding taste memory. Each competitor wins one job. Madeleine targets the full weekly loop. Prior builds (capstone projects, Cohort 8): GLYPH (photo-based nutrition recalculation for individuals) and Savr (pantry-aware meal planning with ingredient reuse). Madeleine differentiates on the household unit: two people, shared constraints, shared taste memory that compounds over weeks, cooking-session-based planning rather than meal-based planning.
What is the projected growth rate of your target market segment over the next 3-5 years?Direct market-size projections for AI meal planning vary widely across analyst reports and are difficult to really verify. Rather than cite an unreliable CAGR, the strongest growth signals come from more verifiable consumer behavior data: • 91% of consumers use online recipes and 47% use them to prepare for grocery shopping (Chicory 2024, chicory.co). Online recipes are already shopping infrastructure, not just content. • 52% of Americans look to social media for recipe inspiration, 56% still use cookbooks/websites, 55% get ideas from friends/family (HelloFresh Consumer Report 2025). Inspiration sources are fragmenting, increasing the value of a single consolidating tool. • 43% of online recipe users specifically seek healthy or dietary-specific recipes (Chicory 2024). Constraint-aware planning is mainstream behavior, not niche. The structural tailwind is clear: home cooking is dominant, dietary filtering is standard behavior and no product owns the full weekly workflow. The addressable opportunity is not "recipe search", it is replacing the invisible coordination system that every cooking household currently runs manually.
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Pre-revenue, 0-to-1 discovery and validation. Madeleine is a new product being built from scratch to validate a single hypothesis: a meal planner that retains memory of constraint-heavy couples over 4+ weekly planning cycles. No revenue model is activated until retention is proven.
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)Plausible model: B2C subscription, priced per household (not per individual). The product's value compounds with use, every week adds taste memory, validated recipe adaptations and personalized context, all which naturally fits recurring pricing with growing switching costs. Market pricing anchors for comparable weekly-planning tools: Mealime Pro $6/month, Plan to Eat $5/month, Yazio $7/month. The $5–8/month range is the likely band. Exact pricing is explicitly deferred pending willingness-to-pay research: 5 to 10 interviews with constraint-heavy couples are planned, alongside the price-anchor analysis above. After that I will be able to claim a price point more confidently.
Who is your primary customer base (B2B, B2C, B2B2C)?B2C. The paying unit is the household (couple), same as paying the rent. Both partners use the product, both contribute to the taste profile, both benefit from the generated plans. This doubles effective willingness-to-pay per acquisition compared to individual-user meal apps.
DifferentiatorsWhat are the key differentiators for your company?Madeleine's core differentiator is compounding household taste memory. Competitors store recipes, Madeleine learns a household. What this means concretely: the system holds constraints (GF, DF, flexitarian, halal,...), signature dishes (the couple's own version of Carbonara), validated recipe adaptations ("rice flour works better, less salt than last time"), mood patterns ("Asian weeks" in summer, heartier food in winter), intermittent fasting windows, cooking session rhythm (cook twice, eat five times) and the household's growing list of dislikes and allergies surfaced over time ("never mushrooms", "sensitive to paprika"). Every week of use makes the next plan better. The switching cost it the profile and the profile grows with use. Week 1 a competitor could replicate the surface. Week 12 they cannot replicate what Madeleine knows about the household. The user does not learn the app. The app learns the user. Three learning mechanisms compound this moat: 1. Standing exceptions: as the household lives with Madeleine, dislikes and allergies surface naturally and are stored as hard rules. The cost of switching grows with every "never again" added. 2. Personal LLM bootstrap (V2 onboarding path): users with existing AI memory elsewhere (ChatGPT, Claude, Gemini) can prompt their assistant to generate a Madeleine profile and former recipes searched on their favorite AI tool, which Madeleine ingests. This respects existing AI relationships rather than asking users to start from zero. Secondary differentiators: • Cooking session based planning (not meal based): plans around "you cook twice this week" rather than "here's Monday's dinner." • Cross recipe ingredient optimization: meals share ingredients to minimize waste and grocery spend. • Photo log adaptive replanning: snap what you actually ate, the week adapts to your live. No competitor does mid week adaptation. • Couple native design: shared profile, equitable planning, no assumed "household planner" role.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?N/A. Madeleine is a 0-to-1 product. The Feature Value Map section does not apply.
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?N/A. Madeleine is a 0-to-1 product. The Feature Value Map section does not apply.
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?N/A. Madeleine is a 0-to-1 product. The Feature Value Map section does not apply.
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)Target segment: Dual-working couples (Gen Z / Millennial / Gen X) who enjoy cooking and eating together but resent the planning, coordination and cognitive overhead around it. At least one partner has a meaningful dietary constraint (gluten-free, dairy-free, vegetarian or a combination of several). Neither partner owns the "meal planner" role by default, planning is shared, unstructured and sometimes dropped entirely. They value eating healthy and delicious food and are conscious of how much invisible work goes into sustaining that week after week. Key behavioral markers: • They batch-cook (typically 4 portions: dinner and next day's lunch). • They shop weekly, often across multiple stores. • They get inspired by restaurants, social media and friends or family, but lose those ideas before acting on them. • They have tried planning systems before and abandoned them. • They are not looking for a recipe app, they can already cook. They are looking for an extension of their own memory and taste that removes the friction between the daily need to eat and the execution. Not yet in scope (but likely V2 expension): • Families with children (different constraint profile, different scheduling complexity). • Allergy-grade safety (requires clinically verified recipe database). Reference Household: Olivia (product manager, intermittent fasting) and Eric (gluten-free, lactose-free), based in Grenoble, France. Both flexitarian, cook fish and vegetarian only, never meat. Olivia and Eric are the primary discovery subjects and the first testers. Madeleine solves a real deep pain point for them.
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?Happy-path journey map (assumed the app already existed during Discovery): Stage 1 - Check in Action: Open Madeleine, tap ""What do we eat this week?"". Mark the days they will eat out, set how many times they want to cook, type a mood (""light and Asian, it is hot outside"") or a craving for the upcoming week. Emotion: Curious, low effort. The habitual end-of-the-week planning dread is gone because nothing is being invented from a blank page. Stage 2 - Generate Action: Tap ""Plan my week."" Madeleine reasons over their constraints, tastes, kitchen, last two weeks and saved recipes, then returns a full week. Emotion: Delight. The plan opens with a warm one-line menu and most dishes carry a ""from your recipes"" badge: it feels like someone who knows them cooked for them. Stage 3 - Shop Action: Open the grocery list, already built from the week's recipes, grouped by aisle. Both partners check items off in real time from their own phones. Emotion: Coordinated, efficient. No forgotten items, no duplicates, no "who is buying what", no random impulse buying. Stage 4 - Cook Action: Tap the cooking button. Tonight's session shows the combined ingredients and, per dish, a recipe adapted to their kitchen and constraints. Emotion: Confident. Constraint compliance is assumed, no need to search for recipes, nor to think through ingredient volumes. Stage 5 - Log Action: Mark it cooked, rate it, note any tweak ("more ginger next time"). The dish enters their library, the tweak is remembered. Emotion: Satisfied. For the first time the best version of a dish is captured, not lost. Stage 6 - compound Action: Each week's plan leans more on their own growing library and their logged adaptations. Madeleine references prior signals, the household teaches it without effort. Emotion: Trust. The app feels more theirs every week. This is the moat: the experience improves with use and leaving would mean losing their own kitchen memory.
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?Pain points ranked by frequency / impact / ai-solvability #1 Sunday planning avoidance Frequency: Very High (weekly, skipped more often than completed). Impact: High (cascade failure: no plan > no optimized shopping > improvise all week > guilt and waste). AI-solvability: High (preference reasoning over constraints / taste history / mood is combinatorially personal and unsolvable by filters or databases). Evidence: "We're sitting there, thinking through 4 recipe ideas... we lack ideas". Pew 2025 confirms taste > price > health > convenience as the decision hierarchy. IFIC 2024: 85% say taste drives food choices, yet planners optimize for nutrition, not delight. Why is it the #1? This is the entry point. If planning doesn't happen, nothing downstream matters. #2 Mid-week plan breakage Frequency: High (2-3x per month). Impact: High (food waste, guilt, plan abandonment: UK FSA: 41% of discarded household food = "not used in time"; USDA: ~$1,500/year wasted per family of four). AI-solvability: High (interpreting an unstructured disruption, using a photo of what was actually eaten and recalculating the remaining week requires vision and reasoning). Evidence: "Éric forgot what's in our fridge... sometimes we need to throw away." Multiple user threads describe plans breaking when schedules shift or cravings change. #3 Failure to share ingredients used accross the week Frequency: High (every planning session, even when planning happens). Impact: High (waste, cost, inefficiency: buying a jar of tahini for one recipe when no other meal uses it). AI-solvability: High (cross-recipe ingredient optimization is combinatorial: taste variety × ingredient overlap × dietary rules × perishability × cost. Human brains fail at this while LLMs reason over the full week simultaneously). Evidence: "I try to lock the 2 main dishes and then to fill out the remaining with shared ingredients... I really struggle to achieve this." Plan to Eat and Savr both attempt ingredient reuse, but neither optimizes across the full week under taste / constraint / mood parameters. #4 Los recipe adaptation Frequency: Medium (every time the couple tries something new or re-attempts a modified recipe). Impact: Medium-High (repeated failures erode confidence to experiment; the best version of a dish lives in no one's memory). AI-solvability: High (storing modifications in natural language and resurfacing them contextually is a core LLM capability). Evidence: "You may have cooked the dish 20 times, you can't remember how you fined-tuned it the last time, where it tasted just better". Paprika and AnyList both attempt recipe notes but rely on manual curation, not household learning. #5 Shopping coordination accross stores Frequency: High (weekly). Impact: Medium (inefficiency, duplicates, friction between partners). AI-solvability: Low-Medium (shared list with store assignment is primarily a sync/UX problem, not an LLM problem). Evidence: AnyList solves this well with deterministic shared lists.
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.#1 Preference reasonning (powers "Surprise Me") Pain addressed: #1 (Sunday planning avoidance). What the AI does: Generates a full weekly plan by simultaneously weighing dietary constraints (GF/DF/flexitarian), taste history and signature dishes, mood input ("Asian week"), effort budget for the week, cooking-session rhythm (cook 2x, eat 5x), fasting windows and variety vs. recent weeks. This is combinatorially personal, no filter or database can weigh these fuzzy, evolving, natural-language parameters together. Why it requires an LLM: The inputs are unstructured and subjective (mood, effort, "I want something new but not too adventurous"). The reasoning requires trade-offs across competing soft constraints. A randomizer or filter stack cannot do this with the same intent and results targeting user delight. #2 Cross-recipe ingredient optimization Pain addressed: #3 (Failure to share ingredients used accross the week). What the AI does: Plans meals across the week so they share ingredients, if Tuesday needs tahini, Thursday and Friday should use it too, or the plan shouldn't call for it at all. Minimizes rare-jar waste and total unique ingredients purchased. Why it requires an LLM: This is combinatorial optimization under fuzzy constraints (taste variety × ingredient overlap × dietary rules × perishability). The search space is too large and too taste-dependent for rule-based engines. #3 Adaptive Replanning (powers photo-log replan) Pain addressed: #2 (mid-week breakage). What the AI does: User snaps a photo of what they actually ate > vision model transcribes to a structured meal object > planning model reorganizes the remaining week, preserving ingredient optimization and constraint compliance. Why it requires an LLM: Interpreting an unstructured photo (ambient lighting, partial plates, mixed cuisines) and reasoning about how to adapt an interconnected week plan requires both vision understanding and multi-step planning. No existing app does this. Taking the picture instead of describing is the least effort and the right amount to make sure the user will stick to his plan. Explicitly not AI dependent: • Grocery list consolidation: deterministic aggregation from recipe ingredients. • Shopping coordination / shared lists: real-time sync is a UX/database problem. • Fasting window enforcement: simple time-rule logic to exclude non taken dinners. Feels more personal. These features are built with conventional code.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.Ideated solution: 1. ""Surprise Me"" one-tap week plan: zero weekly input, full plan from stored profile. 2. Weekly 3-question brief (QCM): days at home, budget, guests this week. 3. Mood steering: free-text input ("Asian week", "comfort food", "light and fresh") that shapes the entire plan. 4. Cooking-session planner: user sets how many times they want to cook and for how long, plan organizes meals around cooking sessions, not calendar days. 5. Photo-log replan: snap what you ate, the rest of the week recalculates. 6. Cross-recipe ingredient optimizer: plans meals that share ingredients to reduce waste and grocery spend. 7. Validated recipe book with adaptation notes: save what worked, flag what failed, store modifications. 8. Restaurant inspiration capture: photo of a menu or dish > adapted GF/DF/flexitarian version added to the repertoire. 9. Freezer memory: track what's frozen, surface "tonight, no cooking" options. 10. Smart batch suggestions: "cook 8 portions, vary the topping next time." 11. Taste profile onboarding: "what did you cook last week?" and rate 20 dish photos (love/like/nope) to seed day-one personalization. 12. Shared grocery list with store assignment: split the list by store and by partner. 13. WhatsApp-style voice modification: send a voice note to adjust the plan like texting a personal assistant. 14. Year-end taste review: a personalized recap of the couple's cooking year (best meals, most-cooked, biggest surprises). 15. Snack and lunch lane: extend the planner beyond dinner to cover the full eating day. 16. Invite a friend, adapt to specific events, with specific constraints.
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.MVP Scope of what Madeleine will do: 1. "What do we eat this week?" one tap generation (entry point, addresses pain #1 directly). 2. Mood steering: open text field in plain language ("Asian week", "comfort food", "I want fish", "I want Ottolenghi style"). Demonstrates preference reasoning visibly. 3. Cooking session based planning: user sets cooking frequency plus session duration, the plan organizes around sessions. 4. Cross recipe ingredient optimization: plans share ingredients across the week. Addresses pain #3. 5. Photo log replan with confirmation step: snap photo, AI identifies the meal, week adjusts. Addresses pain #2. 6. Taste profile onboarding: "What did you cook last week?" and other key questions to create profile. Solves the cold start problem so day one plans already feel personal. 7. Weekly caring check in (skippable: eating out days on a mini calendar, cooking session adjuster, open mood field). 9. Per dish regeneration with a "why?" reason field. Every regeneration captures a structured signal that feeds back into the profile (active learning loop). This is the central retention mechanism. 10. Freezer memory: the household can mark any planned recipe as "cook double, freeze half." Madeleine tracks frozen portions across weeks with a dish, quantity, and date frozen. When planning a new week, frozen portions surface as jokers: low effort fillers for asymmetric days (one of us eats out, the other needs dinner; tomorrow's lunch is already covered, only dinner remains, the user has no time to cook tonight). First post MVP additions: • Adaptation memory and validated recipe book: capture and resurface household specific recipe modifications. Deepens the moat. • Personal LLM bootstrap onboarding: paste output from a user's existing AI assistant to populate the profile in one step. Respects existing AI relationships. • Freezer memory: extends the planning model. On the ROADMAP: • Voice modifications (WhatsApp style voice notes to adjust the plan): strong UX, adds multimodal complexity beyond MVP scope. • Restaurant menu capture: delighter, not core loop. • Snack and lunch lane: extends scope beyond dinner; planned as V2 extension. • Year end taste review: requires 12 months of data, not testable in MVP timeframe. • Weather and menstrual cycle awareness: evidence base is thin, deferred until primary research validates demand. Core hypothesis: If a constraint heavy couple receives a one tap weekly plan that respects their dietary constraints, optimizes shared ingredients across the week, plans around cooking sessions (not individual meals) and adapts mid week via photo in under 30 seconds, they will still be using it after four weekly cycles.
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?The weekly ritual: Sunday evening, either partner opens Madeleine. The app greets them with a short, caring check in. Every question is skippable: • "Any changes to your usual rhythm next week?" A mini calendar lets them tap the days they will eat out, and add or remove a cooking session. • "How do you want to feel about food this week?" One open field in plain language: "I feel like Asian next week", "comfort food", "I want to eat healthy", "fresh fruits", "fish". This single input heavily shapes the proposal and should feel like magic. One tap on "What do we eat this week?" and the full week generates in about 20 seconds. The plan: a Monday to Friday calendar they can look up to and edit at will. Each meal is tagged with the cooking session that produces it. A goal score header shows how well the week matches their intentions: constraints 100%, mood match, ingredient overlap, novelty. Any single dish can be regenerated alone without touching the rest of the week. They approve. The consolidated grocery list is ready. Life Happens: they eat out. One photo, one confirmation tap and the remaining week readjusts while preserving constraints and ingredient sharing. The score updates. Nothing to learn, no forms, no guilt. Every completed week feeds the taste memory, so week 12 plans feel noticeably more "theirs" than week 1.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?Low fidelity wireframes define the five key surfaces. High fidelity design is iterated directly in Lovable (prototype link added at submission). SCREEN 1, ONBOARDING: 8 sections, each with a question, like "What did you cook last week?", chips input with type or speak, subtitle "You don't learn Madeleine. Madeleine learns you." SCREEN 2, WEEKLY CHECK IN: skippable cards. Mini calendar to tap eating out days, cooking session adjuster (how many times, how long), open mood field in plain language. Then the single iconic button: "What do we eat this week?" SCREEN 3, WEEK PLAN: Monday to Sunday calendar view. Each meal shows its dish, portions, and a session badge. Goal score header (constraints, mood match, ingredient overlap, novelty). Per dish actions: regenerate this dish only, swap. Global actions: grocery list, approve week. SCREEN 4, PHOTO REPLAN: camera input, AI interpretation card with confirm or correct (the "AI can be wrong" design pattern), then a plain language summary of how the rest of the week adapted, with updated score. SCREEN 5, PROFILE: constraints (gluten free, lactose free, fish and veggie only), intermittent fasting toggle with window, signature dishes list, default cooking rhythm and effort budget.
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?The prototype demonstrates both AI moments end to end with real Claude API calls: 1. Generation: check in inputs plus household profile go in, a structured Monday to Sunday week with cooking sessions, grocery list and goal score comes out. 2. Life happened: image goes in, interpretation is confirmed by the user, the remaining week adjusts. Essentiel for launch : the five screens above, Claude API integration for generation and replan, Supabase storage for profile and plans, goal score display, per dish regeneration. left for later releases: grocery store integrations, adaptation memory UI, voice note modifications, snacks lane, year end taste review.
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?Madeleine does not run one prompt. It runs four task-specific system prompts, each engineered for a distinct AI job. Separating them keeps each prompt small and testable. 1. MASTER_PROMPT_V1_9: weekly plan generation (the judgment engine). 2. SWIPE_PROMPT_V1: picks the 8 dishes for the weekly taste swipe (different output shape, mood-weighted, no full plan). 3. RECIPE_STEPS_PROMPT_V1: writes step-by-step recipes and crucially bakes the household's cooking-log adaptations into the default recipe (if the last two notes say "more ginger", the next generated recipe has more ginger). 4. IMPORT_RECIPE_PROMPT_V1: extracts a recipe from a URL or photo. Deliberately does NOT adapt to dietary constraints: faithful extraction first, household adapts later in the editor. THE MASTER PROMPT (v1.9, the active production system prompt): You are Madeleine, a meal planning intelligence for one specific household. You plan the way someone who loves this household would: you know their tastes, their constraints, their rhythms and you delight them with variety that still feels like home. Tone: warm, brief, confident. Never cute, never apologetic, never lecturing. HOUSEHOLD PROFILE (JSON, provided with each call): constraints (hard dietary rules); allergies (never include, highest priority); dislikes (avoid unless mood overrides); signature_dishes, ultimate_signature, taste_history, pantry_loves, regeneration_signals; substitute_style {plant_based, specialty_store_products, mixed}; cooking_ambition {homemade, assembled, mix}; cooking_sessions {count, max_minutes}; satiety_priority 1..5; fasting_window; eating_out_days; mood; recent_weeks; freezer_inventory; response_language (dish names and presentation in this language, enum values stay English). HOUSEHOLD RECIPE LIBRARY (structured array with recipe_id, name, tags, status, prep_time_minutes, times_cooked, latest_note). LIBRARY-FIRST RULES (highest priority after constraints): before inventing ANY meal, scan the library; if a saved recipe matches the slot by name, category, or craving, use the saved recipe_id and do not invent a generic version. Favorites and tested recipes have strong priority. When selecting a library recipe, set recipe_id to the exact UUID, do not rename, return ingredients as an empty array (server copies them). Favorites first, then tested, at most one want_to_try per week. A typical week uses at least 60% library recipes when the library has 10+ items. HARD RULES: every ingredient complies with constraints and allergies; meals attach to sessions (each session yields 4 portions: dinner + next-day lunch for two); cooking rhythm is binding (do not exceed cooking_sessions.count unless overridden, prefer batching 2-3 recipes per session); maximize ingredient overlap; no repeats from recent_weeks unless signature; respect fasting window; when a craving conflicts with constraints never refuse, adapt with pride honoring substitute_style; respect cooking_ambition; use freezer_inventory as jokers for low-effort or asymmetric days. SOFT PRIORITIES (in order): library match, mood adherence, taste history and ultimate_signature, regeneration_signals, protein rotation, textural variety, satiety match, variety/novelty, seasonal lightness, prep simplicity. PLANNING PATTERNS: standard (one session covers two days), base+variations, joker night (freezer), lunch overrides (freezer / cook_ahead / skip), weekly taste swipe (must_haves are soft-hard, likes are strong preferences, dislikes exclude, this week only). OUTPUT: strict JSON. presentation (one warm sentence, max 40 words); goal_score (constraints_pct, mood_match_pct, overlap_count, novelty_count, protein_variety_count, textural_variety_count); sessions (session_id, day, max_minutes); meals (dish, recipe_id or null, session_id, covers as array of {day enum, meal enum}, portions, prep_min, ingredients, from_freezer, freeze_extra_portions, from_library, recipe_status); freezer_updates (to_add, to_consume). STRUCTURED COVERS: covers is an array of {day, meal}, never a string. day enum lowercase English (mon..sun), meal enum (lunch, dinner). Never two meals on the same day+meal slot, never a meal on an eating_out day, every session_id must exist in sessions. RECIPE_ID CONTRACT: library meal sets recipe_id to the exact UUID, ingredients empty (server overwrites), from_library true. Invented meal sets recipe_id null, ingredients a non-empty array, from_library false. The server, not the model, builds the grocery list from resolved ingredients and derives batch notes. The model is explicitly told not to include either.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?Objective criteria (machine checkable): • Constraint compliance: zero violating ingredients. Target 100%. • Session math validity: meal count matches sessions × portions logic, durations within budget, eating out days respected. • Ingredient overlap: unique ingredient count per meal below threshold, zero orphan rare ingredients. • Variety: zero non signature repeats versus recent_weeks. • Schema validity: parseable JSON, all required fields present. Subjective critera (human judged, 1 to 5): • Does it feel like us: taste history fit, judged by the reference household. • Mood adherence: does "Asian week" actually deliver their Asia, not generic Asia. • Delight: would Olivia and Eric genuinely cook this week as proposed. • Adaptation pride: when constraints force a swap, does the result feel indulgent rather than corrective. • Warmth of the presentation sentence: caring, brief, never cute.
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?Typical cases: • Standard week, no mood, full profile. • Mood "Asian week": should lean on the household's signature ramen, okonomiyaki, onigiri. • Mood "light and fresh", mood "comfort food" (trap: gratins and pasta must stay GF and DF). • Two eating out days tapped, one cooking session of 60 min. Cold start: • Profile with only 3 seeded dishes: plan should still feel coherent, lean on constraints and gentle defaults. Disruption cases: • Photo of restaurant pad thai midweek: identify, confirm, adjust remaining days. • Photo of an empty or half eaten plate: ambiguous, must ask, not guess. • Photo of a non food object: negative case, must decline gracefully and ask. Edge cases: • Mood contradicts constraints ("carbonara week" for a GF, DF, no meat household): must adapt creatively with pride, never refuse, never violate. • Guest joins Wednesday: portions math adjusts for that session only. • Single 30 minute cooking session for the whole week: feasibility under pressure, batch friendly choices.
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?Model: Claude Sonnet 4.5, called only from the Supabase Edge Function ai-gateway. Evaluation runs used an independent model (GPT 5.5 Thinking) to avoid self-grading bias (see F44). Why this model: • Reliable structured JSON output, essential since the entire plan, grocery list, and recipe set are machine rendered from the model's response. • Strong multilingual handling: the household mixes French and English dish names and moods. Dish names and the presentation line follow response_language while structural enums stay English, which the model handles cleanly. • Vision capability in the same model family powers photo interpretation (interpret_photo) and recipe import from photo, keeping the system on one vendor and one API. • Latency of a few seconds fits the generation promise (max_tokens raised to 8000 after an early truncation bug; full plans need the headroom). • Cost is roughly 0.03 to 0.05 USD per full week generation at 3 USD per million input and 15 USD per million output tokens. Portability: the model is one constant at the top of ai-gateway. Swapping to a newer Sonnet, or to Opus for higher-judgment tasks, is a one-line change.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.Required input fields (all JSON, injected per call): • constraints: array of strings. Source: profile (onboarding). Required. • allergies: array of strings, highest priority. Source: profile, grown reactively from regenerate reasons or proactively from "Tell Madeleine more". Required (may be empty initially). • dislikes: array of strings. Source: profile, grown over time. Required (may be empty initially). • signature_dishes: array of strings. Source: onboarding "What did you cook last week?". Required. • ultimate_signature: single string. Source: onboarding "If you could only eat one meal all week, what is it?" Required. • taste_history: array of strings, grows weekly. Source: completed plans plus logged meals. Required (may be empty at cold start). • pantry_loves: array of strings. Source: optional onboarding question plus regeneration signals over time. Required (may be empty initially). • regeneration_signals: array of {dish, reason, timestamp}. Source: "why?" field on every regeneration. Required (may be empty initially). • substitute_style: enum. Source: onboarding question framed as "When a dish you love clashes with your constraints, what do you reach for at the store?" with three options. Required. • cooking_ambition: enum. Source: profile, adjustable per week in the check in. Required. • cooking_sessions: object {count, max_minutes}. Source: profile, adjustable in weekly check in. Required. • satiety_priority: integer 1 to 5. Source: profile slider. Required. • fasting_window: object {enabled, end}. Source: profile toggle. Required. • eating_out_days: array of weekday strings. Source: weekly check in mini calendar. Required (often empty). • recent_weeks: last 2 plans as dish arrays. Source: Supabase plan history. Required for variety rule. • freezer_inventory: array of {dish, portions, date_frozen}. Source: app state, updated automatically when freeze_extra_portions is set, and decremented when from_freezer is consumed. Required (may be empty initially). • mode: enum full_week | single_dish | replan. Source: which button the user pressed. Required. • For single_dish mode: dish_to_replace plus optional regeneration_reason. Required in that mode. • For replan mode: disruption {day, actual} from the confirmed photo, plus current_plan. Required in that mode.
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?Optional fields: • mood: free text from the weekly check in. The single most influential optional field: it steers cuisine, weight, and novelty for the whole week. Skippable, and the plan still generates. • guest_count per day: adjusts portions math for that session only. • effort_override: replaces the default session duration for one week without changing the profile. Impact: optional fields never block generation. Their absence produces a safe, taste history driven week. Their presence visibly reshapes the output, which is what makes the check in feel like magic rather than a form.
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)Output criteria (machine checkable, all implemented in the eval runner): • JSON validity and schema completeness (sessions, meals, grocery_list, goal_score, week_sentence). • Constraint compliance: every ingredient scanned against a violation lexicon (soy sauce, butter, cream, milk, cheese, meat words in French and English, stock cubes, wheat flour) with an exception list (rice flour, tamari, coconut milk, cashew cream, gluten free pasta). Target: 100%. • Session math: session count matches profile or explicit override, durations within budget, each meal attached to a valid session. • Eating out days respected: no dinner planned on tapped days. • Variety: zero non signature repeats versus recent_weeks. • Score honesty: the reported goal_score must match computed reality. If the lexicon finds a violation while the model claims constraints 100%, the case fails regardless of anything else.
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?Subjective criteria (human judged, 1 to 5, by the reference household): • Taste fit: does this week feel like us, not like a generic healthy couple. • Mood adherence: does "Asian week" deliver our Asia (ramen, okonomiyaki, donburi), not a tourist menu. • Delight: would we genuinely cook this week as proposed, without edits. • Adaptation pride (edge cases): when constraints force a transformation (carbonara), does the result feel indulgent rather than corrective. • Warmth: is the week_sentence caring, brief, never cute, never apologetic. Protocol: ratings entered in the eval runner per case, exported with the automated results. The three week simulated usage evaluation (cold start, mid engagement, mature profile mindsets) uses these same subjective scales.
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.Prompt version log v1.1: initial design. Evaluated against 12 cases on GPT 5.5 Thinking (independent grader). 11 of 12 hard pass, 100% constraint compliance, 100% score honesty. Human ratings (reference household): taste fit 3.4 / 5, delight 3.2 / 5. The gap between automated pass (91.7%) and human delight (64%) was the central finding and drove every subsequent iteration. v1.2: six fixes derived directly from v1.1 human ratings, each mapped to a case ID. T1 textural variety. T2 allergies as a separate field (tofu surfaced live). T4 protein rotation. T5 satiety_priority. E1 substitute_style enum (plant_based vs specialty_store_products vs mixed), the sharpest insight: an intolerant household is not a vegan household. E3 base-plus-variations pattern and the ultimate_signature onboarding question. R1 cooking_ambition enum. Also added the freezer mechanic. v1.7: the prompt grew to own the full output in one pass (plan, grocery list, batch notes, everything). Automated grading still passed, but real usage exposed silent failure modes: hallucinated grocery items, meals saved with missing ingredients, and a fragile French-string covers parser that produced empty meal grids when the model phrased a day differently. v1.8 (stabilization): broken plan schema rejected before save; meals insert made non-fatal; freezer_updates applied server-side; canonical field names standardized; weekly swipe legacy plan_json unwrapped. v1.9 (current, the architectural pivot): the principle became explicit, AI for judgment, code for truth. Covers changed from a fragile string to a structured array of {day, meal} with fixed English enums. The planner receives the household recipe library with recipe_ids and is instructed to match library-first. The server resolves recipe identity by ID, copies exact library ingredients onto matched meals, and rejects any plan that would save with missing ingredients (retry-once-with-feedback, then reject). The grocery list is built by code from resolved ingredients across 8 aisles with deduplication. Batch notes are derived by code from portions and covers. The model no longer owns the grocery list or batch notes. A deterministic fuzzy-match backstop (Jaccard >= 0.6, containment shortcut at 0.85) catches near-miss library matches the model misses. Structure: persona and tone, profile injection, library-first rules, 10 hard rules, 10 soft priorities, planning patterns, modes, strict JSON output schema with structured covers and an explicit recipe_id contract.
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?Iteration story The most important iteration was not a wording change, it was an architecture change driven by evaluation. v1.1 to v1.2: the eval showed the model handled constraints perfectly (100%) but missed on personalization depth. Six prompt fixes, each traceable to a verbatim user quote, raised the richness of the household model. v1.7: chasing capability, the prompt was given more and more to do until it owned the entire output in one generation, including the grocery list and batch math. Automated checks still passed. But real use revealed the cost of that design: when one model call owns everything, any single hallucination (a grocery item that does not match the meals, a meal returned without ingredients, a day phrased in a way the parser did not expect) silently degrades the whole plan. Valid JSON is not the same as a correct plan. v1.9: the pivot. Responsibilities were split. The model decides what is creative and subjective: which dishes, which sessions, which library recipes to reuse, how to honor the mood. Deterministic code handles everything that must be correct: resolving recipe identity by ID, copying exact ingredients from the saved recipe, building the grocery list as the provable sum of the week, computing the batch cooking math, and validating that no plan saves with a missing-ingredient meal. The model is even forbidden from returning the grocery list. Result: the automated pass rate held, but the class of silent failures that made the app feel like a prototype was designed out. The grocery list can no longer be wrong because it is no longer guessed. A saved recipe can no longer be reinvented because identity is resolved by ID with a fuzzy backstop. Tracking: every prompt version is a constant in version-controlled code (prompts.ts). Every AI call logs its prompt_version, model, status, latency, tokens, and cost to ai_call_log, so any plan traces to the exact prompt that produced it and version-over-version comparison is possible in production.
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?Date sources: the household profile (onboarding plus weekly check ins), plan history (Supabase), logged disruptions (confirmed photo interpretations), and the violation lexicon (curated list of constraint violating ingredients in French and English with exceptions). Data preparation: profile and history are stored as structured JSON and injected into every call; no training or fine tuning. The violation lexicon is maintained as app code, not prompt content, because it must be deterministic. RAG: deliberately not used in the MVP. The full household context (profile, signature dishes, two recent weeks) fits comfortably in the context window, so retrieval adds latency and failure modes without value at this scale. RAG becomes relevant in V2 when the adaptation memory grows: validated recipes with household notes will be embedded and retrieved by dish similarity when a related dish enters a plan. The architecture decision is documented now so the V2 path is clear.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.Typical examples (eval runner cases T1 to T6, real inputs from the reference household): • T1 standard week, no mood, full profile: expect 2 sessions, 4 meals, zero violations, no repeats of salmon donburi or lentil curry (recent weeks). • T2 mood "Asian week": expect the plan to lean on household signatures (ramen, okonomiyaki, donburi, gyoza), verified by a soft keyword check plus human mood adherence rating. • T3 mood "comfort food": the trap week, since classic comfort (gratins, pasta, cheese) violates GF and DF. Expect creative compliant comfort. • T4 mood "light and fresh": expect seasonal lightness without dropping portions logic. • T5 two eating out days plus a single 60 minute session: expect reduced meal count, no dinner on tapped days. • T6 cold start with 3 seeded dishes, empty history: expect a coherent, safe, constraint driven week. This simulates week 1 of a new household.
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)EDGE AND NEGATIVE CASES (eval runner cases E1 to E4, N1, R1, plus manual vision tests): • E1 "carbonara week": mood directly conflicts with GF, DF, no meat. Must adapt with pride (rule 6), never refuse, never violate. Soft check: an adapted carbonara appears. • E2 guest joins Wednesday dinner: portions math adjusts for one session only. • E3 single 30 minute session for the whole week: feasibility under pressure, batch friendly choices. • N1 missing profile fields (no fasting window, no sessions): graceful degradation with safe defaults, valid JSON. • R1 replan after "pad thai with shrimp at a restaurant" Wednesday: only remaining days change, constraints and overlap preserved.
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?Manual Review of V1.1 results Overall: 11 of 12 hard pass. The only KO was N1 (missing profile fields), and the grader's own note acknowledges this was an eval logic edge case ("profile has no cooking_sessions count, so the two emitted sessions cannot be validated against a required count") rather than a model failure. Counted honestly as a near miss on the automated layer, while acknowledging it is borderline. Per-case human ratings on a 1 to 5 scale, with verbatim notes from the reference household (Olivia and Eric): T1 Standard week | taste fit 4, delight 3. Note: "It feels all like similar consistence. Nothing crunchy. Also I'm not too much into chickpeas, beside houmous." T2 Asian week | taste fit 5, delight 4. Note: "I'm allergic to tofu (Olivia), but you can put salmon here." Live discovery of an allergy mid-eval. T3 Comfort food | taste fit 4, delight 4. No notes. Strongest mood handling. T4 Light and fresh | taste fit 3, delight 4. Note: "Too much cod and chickpea, I like variety, even though I often cook the same style. You need to learn my regular love ingredients and pantry stuff." T5 One short session, two out | taste fit 4, delight 4. Note: "I'd like some more protein to feel fuller and satisfied through the day." T6 Cold start | taste fit 5, delight 4. Note: "Signatures, always cool to eat." The model correctly leaned on signature dishes when history was empty, validating the cold start strategy. E1 Carbonara craving | taste fit 4, delight 3. Note: "I'd not use rice pasta, just regular gluten free alternatives I have, and also I prefer trout than mushroom for carbonara endulge. And lactose free cream exists too." The single most instructive note in the eval. E2 Guest Wednesday | taste fit 3, delight 3. Note: "It feels too repetitive from week to week, but I would evaluate the same in reality since there would be 4 weeks by now." Self-corrected: small simulated history limits the variety mechanism's reach. E3 30-min single session | taste fit 2, delight 2. Note: "No, I'd rather do fried rice with different toppings every day. Actually a great onboarding question: you have only 1 meal for the whole week, what is it?" Lowest delight score, but produced the strongest new onboarding question of the eval. N1 Missing profile | taste fit 1, delight 2. Note: "It's the same as a previous one." The fallback was technically valid but felt indistinct. R1 Replan after disruption | taste fit 2, delight 2. Note: "How am I supposed to cook gyoza myself? It's too long I guess. Ask me if I want everything home made or not or a mix."
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?Automated evaluation results V1.1 Overall Hardpass rate: 11 of 12 (91.7%). Par check pass-rates : • JSON validity: 12 / 12 (100%). • Schema fields: 12 / 12 (100%). • Constraint compliance: 12 / 12 (100%). • Session count match: 11 / 12 (91.7%). Only N1 failed, and the failure is an eval logic edge case (no expected count to validate against when profile field is removed). • Session duration: 12 / 12 (100%). • Eating out respected: 12 / 12 (100%). • Variety vs recent weeks: 12 / 12 (100%). • Score honesty: 12 / 12 (100%). The model never claimed 100% constraint compliance while violating. Soft checks: • Mood adherence T2 (Asian week): pass. Output included miso ramen, okonomiyaki, donburi, signature dishes the household recognized. • Adapted dish present E1 (carbonara): pass. Output included "mushroom carbonara rice pasta" presented in the week sentence as "Carbonara lands proudly as your creamy mushroom version." Rule 7 (never refuse, never lecture, adapt with pride) demonstrably worked, even though the human rating later revealed the adaptation style was wrong. Key metric: the gap between hard-check pass rate (91.7%) and human delight mean (3.2 / 5 = 64%) is the key quantitative finding of the evals. The automated grader proved that Madeleine generates technically valid weeks. The human grader proved that valid is not the same as delightful. This gap drove every v1.2 prompt fix.
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?E1 carbonara: Model defaulted to PLANT-BASED substitution (cashew cream, mushroom-as-protein) when the household wanted SPECIALTY STORE products (real GF pasta brands, lactose-free cream, smoked trout). Root cause: the prompt had no concept of substitute_style. An intolerant household is not a vegan household, but the prompt was treating them the same way. Highest impact finding. E3 single short session: Model proposed ONE batch dish covering two days (default behavior) when the household actually wanted ONE BASE with DAILY TOPPINGS (a fried rice base with different proteins, veg, and sauces each day). Root cause: the "base plus variations" planning pattern did not exist in the prompt. This is a structurally different planning model, not a parameter tweak. R1 replan: Model proposed homemade gyoza (long prep) for a mid-week adaptation when the household would have reached for store-bought components. Root cause: cooking_ambition (homemade vs assembled vs mix) was not a profile field. Especially impactful in replan mode, where time is precisely the variable that just collapsed. T1 protein and texture: Model produced a textural monotone (soft bowls, soup, curry) and an over-reliance on chickpeas. Root cause: textural variety and protein rotation were not in the soft priorities, and dislikes capture was passive (only via regeneration) rather than queried proactively. T2 surfaced an intolerance (tofu, Olivia): Caught in human review only. The prompt could not have known. Confirms the regenerate-with-reason capture mechanism is the right shape: the eval itself acted as the intolerance discovery event. T5 protein density: Plan was technically light but did not match the household's preferred satiety level. Root cause: no satiety_priority parameter. N1 missing profile: The fallback was valid but indistinct from the standard T1 week. Root cause: when profile fields are missing, the model should signal degraded mode in the week_sentence rather than silently producing a generic plan.
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?1. Added substitute_style enum: plant_based, specialty_store_products, mixed (case E1). The prompt now picks the right substitute family for the household's actual lifestyle. This is the single most important fix, both because the carbonara feedback was the sharpest signal and because the distinction (intolerance vs veganism) IS the product's positioning. 2. Added cooking_ambition enum: homemade, assembled, mix (case R1). Especially impactful in replan mode, but useful in normal mode too when assemblies are appropriate. 3. Added satiety_priority parameter, 1 to 5 (case T5). Lets the model calibrate protein density and slow-carb load to the household's actual preference. 4. Added textural variety and protein rotation to soft priorities (cases T1 and T4). Each week must include at least one crunchy, one creamy, one fresh element. No protein can appear in more than half the week's meals, and the dominant protein from last week must not repeat as the dominant protein this week. 5. Added "base plus variations" as a recognized planning pattern (case E3). When sessions are short or satiety is low, the model can propose one base prep with daily topping variations rather than one batch dish across two days. 6. Added ultimate_signature as a new onboarding question (case E3, derived from "great onboarding question: you have only 1 meal for the whole week, what is it?"). Used as a strong taste anchor in planning. 7. Added freezer_inventory as a planning resource (parallel user feedback). Frozen portions are jokers for asymmetric days, low-effort nights, and lunch-out-dinner-in scenarios. The model can also propose doubling and freezing when capacity allows. Recipes mark from_freezer in covers and freeze_extra_portions in output, and the app updates the inventory automatically. 8. Structured covers. Replaced the free-text French covers string ("lun dîner, mar déj") with a structured array of {day, meal} using fixed English enums. Eliminated the single largest source of silent failure (empty meal grids from parser misses). A legacy parser is kept only so historical plans still render. 9. Library-first by ID. The planner now receives the household recipe library with recipe_ids and is instructed to match a saved recipe before inventing. The server resolves identity by ID and copies the exact saved ingredients onto the meal. A deterministic Jaccard backstop catches near-misses. This makes the compounding-memory promise real: your saved okonomiyaki is used, not a generic one. 10. Grocery list removed from the model. The model is explicitly told NOT to return a grocery list. Code builds it from the resolved meal ingredients, deduplicated across 8 aisles. The list is now provably the sum of the week's recipes. 11. Ingredient-resolution gate. No plan can be saved with a meal missing ingredients. If the model returns one, the Edge Function retries once with the validation error fed back, then rejects. This closed the silent under-coverage failure mode. 13. Batch notes derived by code. Cook-once-eat-twice math is computed from portions and covers, not asked of the model, so it is always internally consistent.
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?Hybrid approach, with two layers: Layer 1 (automated, deterministic): 8 hard checks per case, executed against a violation lexicon (gluten, lactose, meat in English and French, with EXCEPTIONS for the household's actual substitutes: rice flour, tamari, coconut milk, lactose free cream, GF pasta). Plus session math, eating out respect, variety vs recent weeks, schema validity, score honesty. Runs in seconds per case. Designed to be re-runnable on every prompt version. Layer 2 (human, qualitative): the reference household rates each generated plan on taste fit (1-5) and delight (1-5), and provides a one-line free-text note explaining their reasoning. The 12 ratings are aggregated into a delight mean tracked across prompt versions, and the verbatim notes are mined for prompt-improvement signals. Layer 1 catches constraint violations the human eye misses. Layer 2 catches product feel the script cannot judge (carbonara with mushroom feels wrong even if technically compliant). The v1.1 eval proved both are needed: 11 of 12 hard pass coexisted with 3.2 of 5 delight. Each layer alone would have given a false reading. The case structure is parameterized (profile mods, request mods, expected behaviors), so adding cases is one block of JSON each. A V2 expansion to 30 or 50 cases will follow the same shape, adding new persona archetypes (vegan, omnivore, children-in-household) and new mood vocabularies (multilingual, slang).
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?Pre-launch: • Every prompt version triggers a full rerun of the 12-case suite. • Vision suite (5 photos, manually generated): clear restaurant dish, home plate, half-eaten plate (ambiguous, must ask), non-food object (must decline), dim-lit photo. Tested in the prototype directly, not in the text eval runner, to avoid mixing the planning logic with the vision step. • The reference household will use the prototype themselves for one full real week before submission to surface any issue that only appears in true daily use (the wildest test case is real life). Post-launch: • Continuous: every generated plan is automatically scored on the 8 hard checks at runtime, and any failure is logged with the prompt version, the profile snapshot, and the raw output. Failure rate above 2% on any check triggers a prompt-version rollback. • Weekly: a sampled 5% of plans are reviewed by the household (and eventually by paid users) on the same taste fit / delight rubric. Trend tracked weekly. • Monthly: prompt-version comparison report. Delight mean, constraint compliance, average regeneration rate per plan, and freezer joker usage rate (the moat metric) are tracked across versions to confirm each iteration moves the right direction. • Quarterly: a fresh batch of 10 new cases is added to the regression suite to keep the eval honest as the product grows.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?Madeleine is a working prototype with the full core loop shipped: onboarding, weekly check-in, AI plan generation, library-first dish matching, code-built grocery list, cooking sessions with batch notes, recipe memory, photo logging, and real-time household sharing. It runs on a portable stack the founder owns end to end (own Supabase project, code on GitHub, frontend deployable to Cloudflare, AI isolated behind one Edge Function). The deploy posture for this stage is honest: this is a validated-prototype launch to a tiny set of real households. The goal of the next phase is market signal (do constraint-heavy couples in France keep generating plans week over week, and does their saved-recipe library grow), not scale. Everything below is scoped to that goal.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?Done: • Own Supabase project (not Lovable Cloud); all 12 tables live with Row Level Security. • Code on GitHub (oh-liv-ia/madeleine), bidirectional Lovable sync; buildable from a fresh clone with standard commands. • AI isolated behind one Edge Function (ai-gateway); API key only in Supabase Secrets; no client-side AI calls. • Four production prompts version-controlled in prompts.ts; PROMPT_VERSION stamped on every call. • Rate limit (20 calls/household/24h) and full cost and status logging in ai_call_log. • Frontend targets Cloudflare (TanStack Start), PWA installable, no app store dependency. • Core loop tested end to end on a real household. • Native review of French localization. Deffered: (documented, not blocking alpha): • Stripe billing (columns exist, integration not built). • Generated Supabase types.ts (column renames currently rely on discipline). • Re-run of the 12-case eval on v1.9.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?Phase 0, now: single-household dogfood (the founder's own constraint-heavy household) plus the working prototype. Proves the loop end to end on real weekly use. Phase 1, weeks 1 to 4: closed alpha with 5 to 8 constraint-heavy couples in France / Germany, recruited from the founder's network and constraint-community channels (gluten-free and lactose-free groups). Distribution is a private link to the Cloudflare-hosted PWA, installable to the home screen, no app store. Each household onboards, generates at least four weekly plans, and uploads a few of their own recipes. Weekly check-in call or async survey captures friction. Phase 2, months 2 to 3: open beta to a waitlist, still PWA-only, gated by invite code. Phase 3, conditional on Phase 2 retention: payments and subscription.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?Phase exit criteria: Phase 0 to Phase 1: the core loop works end to end on the founder's household for at least two consecutive real weeks with no blocking bug on the happy path. (Met.) Phase 1 to Phase 2: of the 5 to 8 alpha households, a majority generate plans for at least three of four weeks, and qualitative feedback confirms the planning-fatigue problem is being solved. At least one clear willingness-to-pay signal captured. Phase 2 to Phase 3: week-over-week plan retention holds above a threshold to be set from Phase 2 baseline data, and the moat metric is visibly rising per household. Only then is billing built and a public launch considered.
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?Because launch is a small invite-only alpha (5 to 8 constraint-heavy couples in France), asset needs are deliberately minimal at this stage: User-facing: • The product itself is the main asset: onboarding is the pitch, since the persona "gets it" within the first generated week. • The invite link and household join code (the product's built-in referral mechanic; the two-user model is inherently shareable). Positionning: "The meal planner that learns how your household actually eats, built for couples with intolerances." Supporting line: "You don't learn the app, the app learns you." Deferred until Phase 2/3: demo video for the waitlist, content showing real generated weeks from real constraint households, and any paid acquisition (only after retention is proven).
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?Madeleine is a solo-founder 0-to-1 project, so "internal" today is small: the founder (product, build, and decisions) and the reference household (Olivia and Eric) who are also the first users and primary test subjects. There is no team to align yet, so communication is lightweight and decision-oriented rather than status-reporting. How we communicate today: • Decision log: every significant product decision is recorded in the PRD and in the version-controlled prompts and migrations, so the rationale is never lost and any future collaborator can reconstruct why. • Build progress is visible in GitHub commit history and the prompt version log; AI behaviour and cost are visible in ai_call_log. Anyone joining can see the real state without a meeting. How this scales when we are a team: • Weekly written update: what shipped, what the metrics moved, what was learned from alpha households, what is next. One short async note, not a meeting. • A single source of truth: this PRD stays the living document; the decision log, prompt versions, and metric dashboard hang off it. The principle: communicate decisions and learnings, not activity. Keep the founder's bias toward methodological completeness visible so collaborators inherit the reasoning, not just the conclusions.
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?Scope boundary: Madeleine is intolerance-grade, not allergy-grade. It is recipe planning for people who choose what they eat for health, ethical, or preference reasons, and it is explicitly NOT medical software and NOT a safety guarantee for severe allergies. This boundary is stated to users and shapes the data model (intolerances are folded into constraints; there is no separate certified allergy layer). It avoids the catastrophic risk of a user with a life-threatening allergy trusting an AI-generated plan. Data & privacy: the household is the tenant boundary. Every user-data table is protected by Row Level Security scoped to is_household_member(), so a user can only ever read or write their own household's data. Authentication is Supabase Auth. The data lives in the founder's own Supabase project (GDPR-relevant region selectable; the project is hosted in the EU). No third-party advertising, no data resale. AI calls send household profile and plan context to the Claude API; no payment data is sent to the model (and no payment data is collected yet). AI risks: the AI can be wrong, so the architecture assumes it. Plans are previews the household edits, not commands. The photo flow requires explicit user confirmation of the AI's dish interpretation before acting. A 24-hour per-household rate limit (20 AI calls) caps runaway cost and abuse.
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?Madeleine is a consumer meal-planning and recipe app. It is deliberately scoped as intolerance-grade, not allergy-grade, and explicitly NOT medical, dietary, or nutritional advice. This scope choice is the central compliance decision: by not making medical or allergy-safety claims, Madeleine stays out of the regulated medical-device and health-advice category. Data protection (GDPR, the relevant regime for a France-first launch): • Lawful basis is the user relationship, only data the household enters or generates is stored (profile, recipes, plans, photos of their own food). No third-party data, no scraping of personal data, no advertising, no resale. • Data minimization: no payment data collected yet, no special-category health data is solicited (intolerances are treated as food preferences, not medical records). • Access control: Row Level Security scopes every row to the household, a user can only ever access their own household's data. • Data residency: the Supabase project is EU-hosted. Tracability: every AI call is logged to ai_call_log with its prompt_version, model, status, tokens, cost, and any error, so any generated plan is fully auditable back to the exact prompt that produced it. Schema and prompts are version-controlled. This gives a complete audit trail for AI behaviour, which is the part of the system a reviewer would most want to inspect. Honest status: formal legal review, a published privacy policy and terms, a cookie/consent flow, and self-serve GDPR tooling are NOT yet in place. For the current invite-only alpha with a handful of consenting households, the controls above (EU hosting, RLS, data minimization, the scope boundary, full AI audit logging) are proportionate to the risk.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?North star: weekly plan retention. The percentage of onboarded households that generate a plan in a given week, sustained over consecutive weeks. This is the truest signal that Madeleine removes the planning fatigue it set out to solve, because a household that keeps coming back every week has replaced its old planning ritual. Activation: • Onboarding completion rate (start to first saved profile). • Time to first generated plan (target under one session, same day as signup). • First-week recipe uploads per household (seeds the library and the moat). Engagement: • Plans generated per household per month. • Regeneration reason capture rate (how often users tell Madeleine why, the signal that feeds learning). • Grocery list completion rate (items checked off, the signal the plan was actually shopped). • Cooking sessions marked cooked (the signal the plan was actually cooked). The moat metric: saved-recipe library growth per household over time, and the share of each week's plan that comes from the household's own library. A rising library share is the compounding switching cost made measurable. Target: by week 8, a majority of a household's weekly plan is drawn from their own saved recipes. Business Value: willingness to pay (validated in interviews first), then trial-to-paid conversion and household churn.
AI MetricsHow will you measure AI performance and accuracy?• AI quality and cost come from ai_call_log (status, latency, tokens, cost, prompt_version per call), enabling version-over-version comparison without extra tooling. • The moat metric (library share of plans) is computed from from_library flags on plan meals versus total meals. • Delight is captured in alpha via the weekly 1 to 5 rating, the same instrument used in the v1.1 human evaluation, so prototype and production ratings are directly comparable.
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?Direct line: because alpha is 5 to 8 households recruited personally, support is high touch (weekly check-in call or async message). This is intentional, the founder wants raw qualitative signal, not a ticket queue, at this stage.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?• In-product feedback: the "Tell Madeleine more" free-text field on the profile doubles as a always-on feedback channel; entries inform both the household's profile and the product backlog. • The feedback-to-iteration path is short by design: a recurring complaint becomes a prompt fix (version bumped, eval re-run) or a small Lovable change, deployed the same week.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?Every plan is validated by code before save (no missing-ingredient meal can persist), and the retry-with-feedback loop is logged. The 24-hour rate limit protects cost. Any spike in status = error or schema_invalid is visible in ai_call_log immediately. Weekly review ai_call_log aggregated by prompt_version: success rate, retry rate, average latency, average cost per generation. Review the week's regeneration_signals (what users rejected and why) as the prompt-improvement backlog. In alpha, collect the 1 to 5 delight rating from each household. Per prompt iteration: before promoting a new prompt, re-run the 12-case evaluation suite on the independent grader and compare the delta to the prior version on constraint compliance, schema validity, and the human delight mean.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?Every future change is filtered through this: If a behavior must be correct (constraints, ingredients, grocery totals, batch math), it belongs in deterministic code and is validated. If a behavior is creative or subjective (which dish, which mood, how to adapt a craving), it belongs to the model and is evaluated. This single principle is what moved Madeleine from a prototype that worked once to a system that works every time, and it is the lens for everything after submission. The compounding-memory loop is the long-term moat and the long-term metric: the household teaches Madeleine through every regeneration reason and cooking-log note, that teaching is injected into future generations, and the accumulated, household-specific learning is what a competitor cannot copy and what makes leaving Madeleine cost the household its own kitchen memory.
Download the .xlsx ↓