Food
Sammy
Sammy is an intelligent cooking engine that converts messy online recipe content into a clean, kitchen-friendly cooking workflow. It extracts recipe details, implied equipment, chef tips, and in-line ingredient quantities so users can cook from a blog recipe without the clutter and physical awkwardness of typical recipe sites. The product is designed around the real cooking environment, including legible cook mode layouts and strong guardrails against ingredient or measurement hallucinations.
The problem
Cooking from an online recipe is a fight against the medium. Home cooks scroll past 2,000 words of narrative and ad banners to find the ingredient list, then "pogo-stick" — scrolling up to check amounts and back down to instructions — while the screen dims mid-sauté and messy hands make tapping a hazard. The result is high cognitive load, buried information density, and physical friction in a high-heat, high-moisture environment. Legacy platforms like AllRecipes and Food Network own discovery but fail execution; recipe managers like Paprika and Mela store static text without semantic execution; and generic AI kitchen assistants hallucinate recipes outright — a hallucinated oven temperature can ruin a dish or start a fire.
The solution
Sammy is an intelligent cooking engine that refactors messy blog recipes into a clean, kitchen-friendly workflow. It extracts recipe details, implied equipment, and chef's tips, and injects ingredient quantities inline into the instruction steps so cooks never have to scroll back to the ingredient list. Sammy is integrity-first: it is the Translator, not the Aggregator, tuned for zero-loss extraction that stays true to the user's chosen recipe rather than inventing a new one. Output renders as high-contrast, landscape-optimized cook-mode cards grouped by recipe component. The business model is a freemium micro-SaaS with a deliberately low-friction price point, positioned as an impulse utility.
How it works
The ingestion pipeline pairs deterministic parsing with LLM reasoning. Firecrawl scrapes and strips website HTML and ads, delivering clean Markdown that reduces token noise before the model sees it; a raw copy-paste fallback bypasses paywalls and firewalls. The target production model is GPT-4o-mini with strict JSON mode, chosen for cost (~97% cheaper than GPT-4o), sub-second latency, and deterministic structure. A master prompt with nine behavioral rules enforces zero hallucination on ingredients and metrics, ≥90% tool-inference accuracy, inline measurement refactoring, a culinary density matrix for volume-to-weight conversion, and verbatim chef's-tip extraction that discards blogger fluff. A Screen 3 human-in-the-loop Verification Canvas lets cooks catch any parsing anomaly. Evaluation uses a 15–25 recipe golden set with a hand-built ground-truth JSON target per case; early tests hit 100% on fraction-format and verbatim-tip preservation, with an 85% target for isolating paragraph-buried warnings.
Who it's for
Sammy targets the Everyday Home Cook — the millions of people who execute digital recipes on mobile devices and suffer the physical-digital disconnect. They value kitchen efficiency, waste reduction, and a focused, ad-free experience. The business is B2C moving to B2B2C. Initial focus is direct-to-consumer utility; a future "Send to Sammy" execution API would let food bloggers keep their discovery-phase SEO and ad revenue while Sammy handles the execution phase — turning potential competitors into a distribution channel, akin to Pinterest's Pin button.
Why it matters
The recipe app market is projected to grow at roughly 10.5–12% CAGR through 2030, and ad-fatigue with bloated recipe blogs creates a clear opening for clean execution utilities. By breaking recipes into atomic units — tools, temperatures, timings, quantities — Sammy owns structured data that enables future capabilities competitors can't match, from nutritional scaling to smart-appliance control. The near-term plan is a tightly controlled 4-week closed beta with 10 friends-and-family cooks, run through weekly evaluative gates on ingestion fidelity, tool-inference utility, and ergonomic adherence. Production scales through gated volume expansion — 100 users, then 1,000 — with a migration off the Lovable prototype to an independent stack and LangSmith telemetry. Trust is the brand: in the kitchen, Sammy is the reliable sous chef, not the experimental poet.
The workflow
The PRD
| PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0 | ||||||
|---|---|---|---|---|---|---|
| Your Name: | Samantha Park | |||||
| Your Product: | Sammy | |||||
| Your Industry: | Consumer Food Tech - home prep, personal productivity | |||||
| Date: | May 8, 2026 | |||||
| 4D Method | AI PRD | Instructor Feedback | ||||
| Phase | Activity | Theme | Topic | Key Question(s) | Your Response Include external links to visuals/prototypes as required. | |
| DISCOVERY | Understand your market, business, product & user context | Business Value Map | Market Attractiveness | What industry is your business in? (ie Financial services, Healthcare, Education, etc)? | Food Tech - Consumer | Samantha, the competitive landscape mapping here is the strongest part of the PRD. Splitting competitors into aggregators, managers, and generative wrappers, then positioning Sammy as a "Translator" against each category, gives you a real strategic frame to build on. The pain points are anchored in physical reality (messy hands, screen dimming, scroll fatigue), and the diverge section shows genuine creative range before you narrow. Places to sharpen. First, "Everyday Intentional Cook" is not yet a persona you could recruit for a usability test. How many nights a week do they cook? Do they own a tablet or prop a phone against a jar? Are they following Bon Appetit links or TikTok saves? Adding two or three behavioral anchors would force your prompt design, your UX decisions, and your pricing to get more honest. Second, the Forrester and Gartner references and the "Blended CAGR of 19.9%" need source links or at minimum report titles. Unsourced authority claims weaken a section that is otherwise doing careful work. Third, and this one matters most for Design: your PRD does not yet isolate where deterministic parsing ends and LLM reasoning begins. Many recipe blogs already expose structured data via JSON-LD schema. Extracting ingredients and temperatures from that schema does not require a language model. The genuinely hard problems (inferring implied tools, generating a "Step 0" prep block, handling freeform narrative with no schema) do. Name that boundary explicitly, because it is the difference between a product that needs AI and one that is decorated with it. What does the target workflow look like when you sit down to write Design? |
| What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors? | Headwinds/Challenges: 1. Ad-Dependency of Sources: Most recipe creators rely on heavy ad-insertion for revenue; blocking or refactoring that content can create tension with the source ecosystem. 2. High LLM Latency/Cost: Complex "Semantic Refactoring" requires high-level reasoning which can be slow and expensive at scale. 3. Physical Environment Risks: High-heat and high-moisture environments are inherently "anti-hardware," making tablet/phone usage a constant friction point for users. Tailwinds/Opportunities: 1. The "Ad-Fatigue" Movement: Users are at a breaking point with ad-heavy, narrative-bloated blogs, creating a massive opening for "Clean, simplified UI" utilities. 2. Advances in Semantic Parsing: Rapid improvements in LLM capability (like GPT-4o-mini) now allow for real-time refactoring that is just now possible. 3. Rise of "Kitchen-as-Hub": The post-pandemic trend of high-frequency home cooking has created a more tech-savvy, "prosumer" home cook who demands better tools. Key Competitors: 1. Legacy Discovery Platforms: (AllRecipes, Food Network, Pinterest). They own the "Discovery" phase but fail the "Execution" phase. 2. Recipe Managers: (Paprika, Mela, Pestle). Excellent for storage, but rely on static text. They lack the Semantic Execution (inline data injection) that Sammy provides. 3. Generative AI "Kitchen Assistants": Generic LLM wrappers that "hallucinate" recipes. Sammy differentiates by maintaining the Integrity of the Source while refactoring the UX of the Task. | |||||
| What is the projected growth rate of your target market segment over the next 3-5 years? | Recipe and Meal Planning App market: The global Recipe App market is projected to grow at a CAGR (Compound Annual Growth Rate) of approximately 10.5% to 12.0% through 2030. (Source: Research and Markets) (Source: Straits Research) (Source: Grand View) Wider Market Ecosystem & Smart Home Intersection: Rather than acting as a static content repository, Sammy operates at the intersection of digital meal planning and contextual consumer software workflows. This sub-segment of consumer applications is experiencing heightened velocity due to the massive consumer shift away from passive browsing toward proactive, distraction-free execution in physical environments. By anchoring its value proposition strictly to the "Zero-Friction" kitchen loop—solving real-world physical boundaries like scroll fatigue and small-screen ergonomics—Sammy is positioned to capture high-margin premium subscribers within a rapidly expanding addressable market of tech-enabled home cooks | |||||
| Business Model | What growth stage is your business currently in (e.g., startup, scale-up, mature)? | Pre-Seed / Problem-Solution Fit. We are currently in the 0-to-1 phase, focusing on validating the core "Refactoring Engine" and "Landscape Cook Mode" as the primary value drivers before scaling to micro-influencer and blogger partnerships. | ||||
| How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?) | Freemium Micro-SaaS. Free access for the first 5 saved recipes (The Hook). A low-friction $5/month or $50/year subscription for unlimited storage, advanced ingredient search, and contextual nutrition data. This model prioritizes mass adoption and long-term retention over high per-user extraction. | |||||
| Who is your primary customer base (B2B, B2C, B2B2C)? | B2C moving to B2B2C. Initial focus is direct-to-consumer utility. Growth phase utilizes a B2B2C play via the "Send to Sammy" API for food bloggers, making Sammy the standard execution standard for third-party recipe content. | |||||
| Differentiators | What are the key differentiators for your company? | Differentiators --> Moats: 1. The "Semantic Infrastructure" Moat Most competitors are Aggregators (they collect recipe links). Sammy is a Translator. The Differentiator: While Paprika or Mela show you a cleaned-up version of the original text, Sammy’s proprietary AI engine performs "Semantic Refactoring." Moat: I own the structured data. Because I've broken down recipes into atomic units (tools, temperatures, timings, and quantities), I can do things competitors can’t (i.e. in the future, perform instant nutritional scaling, suggesting how to use up leftover ingredients, or controlling smart appliances.) 2. The new Home Kitchen UX Standard I am the only company building for the physical-digital Kitchen Environment. The Differentiator: "Context-Aware UX." Other apps are designed for "Portrait/Read" mode with scrolling. Sammy is designed for "Landscape + Execute" mode. Business Impact: I am defining the ergonomics of digital-home cooking. By focusing on distance-readability (the "4-foot glance") and inline measurement injection, Sammy becomes the "standard" for how professionals (future) and serious home cooks interact with digital content while their hands are busy. 3. Integrity-First AI (Anti-Hallucination) In a world of "Generative AI" that invents recipes, Sammy is true to the source (Verification Engine). The Differentiator: Sammy maintains the Source of Truth. The AI is tuned for "Zero-Loss Extraction." We don't hallucinate a new way to make lasagna; we perfect the delivery of the user’s chosen recipe. Business Impact: Trust. In the kitchen, a "hallucinated" oven temperature can lead to a fire or a ruined $50 roast. Sammy’s brand is built on being the "Reliable Sous Chef," not the "Experimental Poet." 4. The "Impulse Utility" Revenue Model (TBD) Not chasing the $15/month "SaaS subscription fatigue." The Differentiator: The $5/Year "No-Brainer" Price Point. Business Impact: This creates a massive Customer Acquisition Cost (CAC) to Life-Time Value (LTV) advantage. At $5 a year, Sammy is an "impulse utility." This allows for viral, low-friction growth where the product is essentially "free enough to try, cheap enough to keep." 5. B2B2C Ecosystem Connectivity (The "Sammy Button") (Future phases) Not working against the food bloggers but empowering them. The Differentiator: Instead of stealing traffic, we will offer an Execution API/javascript snippet. No one else has this yet. (think Pintest's Pin button) Business Impact: Bloggers can add a "Send to Sammy" button to their sites. They keep the SEO/Ad revenue from the "Discovery" phase, but Sammy takes over for the "Execution" phase. This turns our biggest potential enemies into your primary distribution channel. | ||||
| Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.) | Customers | Who are the customers (ie buyers) of your product? | [Insert your response here] | |||
| End Users | Who are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context? | [Insert your response here] | ||||
| Current Products / Services | If you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers? | [Insert your response here] | ||||
| User Value Map | Target Persona | Who is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc) | The "Everyday Home Cook." This segment covers the millions of home cooks who use mobile devices to execute digital recipes but suffer from the "Physical-Digital Disconnect." We are targeting users who value kitchen efficiency, waste reduction, and a focused, ad-free cooking experience. Sammy: User Persona, Everyday Home Cook | |||
| Journey Map (current-state) | What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product? | 1. Discovery: User finds a recipe on a high-ad mobile blog. 2. Review: User scrolls past 2,000 words of narrative to find the ingredient list, dismissing ads and videos. 3. Execution: User begins cooking, repeatedly scrolling up to check amounts and back down to instructions ("Pogo-sticking"). 4. Friction: User touches screen with messy hands; constant rechecking ingredient amounts; screen dims/locks mid-sauté; crucial steps are missed due to visual clutter. | ||||
| Pain-points | Where does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe? | 1. High Cognitive Load: Constant mental mapping of separated ingredients to instructions. 2. Information Density: Essential data is buried under ad banners and SEO filler. 3. Physical Friction: Managing a delicate device in a high-heat, messy environment, tapping/unlocking/scrolling to see the instructions and ingredients. 4. "Mustard Finger" UX: The hygiene and hardware risk of touching phones during prep. | ||||
| AI Opportunities | From your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first. | 1 & 2 (above) --> Semantic Transformation: Using LLMs to "read" unstructured text and reconstruct it as an augmented instruction layer (refactored recipe card instructions, with only essential information (no ads!) 1 (above). Inferential Logic: Identifying required tools that are implied but not explicitly listed in a "Gear" section. (i.e. Tablespoon, teaspoon, pot with lid) 1 (above). Procedural Mutation: Injecting "Step 0/Prep" instructions for complex steps or pre-work, in the correct order. | ||||
| Develop an AI Solution Hypothesis | AI Solution Hypothesis | Diverge | Ideate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage. | The "Recipe-to-SMS" Bot: A bot that texts you every ingredient and step one by one. You text back "Done" to get the next one. The "Recipe-to-Spotify" Podcast Generator: An AI that converts any recipe into a 20-minute audio track with ambient "cooking music" and timed instructions. The "Audio-Only" Sous Chef: A pure voice assistant (like an Alexa Skill) that dictates the recipe step-by-step. The "Screen-Lock Disabler" Overlay: A simple browser extension that just keeps the screen from dimming on recipe sites. Sammy (Semantic Execution Layer) Use LLMs to ingest the "spaghetti code" of a blog and re-render it into a high-contrast, landscape-optimized dashboard with inline data injection. The "Dynamic Instruction-to-Timer" Engine Instead of just showing text, this AI analyzes the recipe to build a Linear Timeline. How it works: The LLM identifies every time-sensitive verb (sear, simmer, rest). It generates a "Project Management" style Gantt chart for the kitchen. The "Context-Aware Viewport" (Smart Zoom) This solution focuses on the Visual Hardware of the phone rather than the text itself. How it works: The AI uses the front-facing camera to track the user's distance and gaze. The "Semantic Cross-Reference" Overlay This is a Browser-First approach (like a Chrome Extension for mobile). How it works: It doesn't move you to a new app. It injects "Smart Tooltips" into the existing blog. | ||
| Converge | Rank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project. | WINNER - Sammy (Semantic Execution Layer) Use LLMs to ingest the "spaghetti code" of a blog and re-render it into a high-contrast, landscape-optimized workflow with inline data injection. 2. The "Semantic Cross-Reference" Overlay This is a Browser-First approach (like a Chrome Extension for mobile). How it works: It doesn't move you to a new app. It injects "Smart Tooltips" into the existing blog. 3. The "Context-Aware Viewport" (Smart Zoom) This solution focuses on the Visual Hardware of the phone rather than the text itself. How it works: The AI uses the front-facing camera to track the user's distance and gaze. | ||||
| DESIGN | Define Target State Workflow | UX Flows & Wireframes Suggested Tool: Excalidraw | Workflow (future) | Assuming your product or feature works as desired, what is the target state workflow? | https://miro.com/app/board/uXjVHSDaZ_k=/?share_link_id=35160831101 | Samantha — the Evaluation Criteria section is doing more work than most PRDs at this stage; eight criteria, each with an explicit benchmark and a Definition of Good. That's the muscle this capstone is trying to build. Two pushes: 1. The prompt isn't yet doing Sammy's headline behavior — inline measurement refactoring. Your wireframes promise it, Eval #6 demands 100% precision on it, but Rule 4 of your prompt says "retain… exactly as written" — preservation, not injection. Check with Roger on this; that's the gap to close. and 2, Eight benchmarks need a real eval set. Three example cases can't carry "≥85% recall on chef's tips" or "98% density conversion." Lock 15–25 labeled recipes spanning your boundary cases — British-metric, Schema-less blog, multi-component bake, no-tips, expired URL — and write the expected output for each. That set is the harness Develop will grade against, and it's where you'll find "chef's tip" needs a sharper boundary ("always use Kerrygold" — tip or fluff?). Define it in the set before the prompt has to guess. Roger will go further on the prompt side when he reviews. |
| Build Wireframes | Wireframes | How will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features? | See Balsamiq wireframe prototype + blue stickies for detailed answers. https://balsamiq.cloud/s2pivxr/ph8pgfl | |||
| Develop Prototype to showcase AI interactions | Prototype Screens Suggested Tool: lovable.dev | What aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases? | PRODUCTION MODEL & ARCHITECTURE: All core prompt engineering configurations, including the 9 critical behavioral rules (Rules 1-9) and the 25-recipe evaluation dataset, are engineered and validated strictly against the GPT-4o ecosystem (Target Production Model: GPT-4o-mini) to guarantee zero-hallucination thresholds and fraction parsing integrity. Note on Master Prompt Isolation: The live Lovable prototype functions primarily as a frontend ergonomic layout mock to test user-facing card mechanics and distance-mode interaction heuristics. Lovable controls the prompt during the pilot phase (gemini-2.5-flash). The production grade underlying prompt engineering IP for Sammy is fully decoupled from the prototype interface, ensuring absolute control over model logic prior to final production API deployment. The prototype demonstration will focus on the core value prop (refactored, inline ingredient amounts inside the instructions) and post pone "Prep Mode" features common in competiting apps (Substitutes, Allergy Alerts, Shopping LIsts) Initial input of the recipe URL: AI Input: A raw recipe URL copied to the clipboard. Visually, this is represented by a single-tap "Tap to Paste Recipe Link" action card that reads the user's clipboard data without requiring text selection or manual typing. AI Processing: A dedicated, multi-step Progress Ticker container replaces generic loading spinners. It explicitly details the background ETL pipeline execution state to build user trust: 📥 Scraping website content... 🤖 AI analyzing inferred tools... 🥞 Separating components... AI Output: The Verification Canvas. The messy, raw blog text is output as clean, structured, modular visual cards that automatically group ingredients and steps by recipe components (e.g., The Cake vs. The Frosting). AI Inferred Tools are explicitly output as isolated high-contrast text tags ([Sifter], [Stand Mixer]) inside each section so cooks can gather missing gear upfront. Essential for Launch: - Deterministic + LLM Ingestion Pipeline: High-fidelity URL parsing fallback mechanism (extracting raw recipe strings when clean Schema.org JSON-LD structured data is missing from a blog). - Automated Semantic Component Grouping: Automatically dividing messy recipes into clear, card-based sub-sections (e.g., separating crust layers from fillings). - Basic Tool Inference Engine: Extracting implicit gear requirements (e.g., knowing that "sift the flour" implies a sifter is needed in the prep checklist). - Manual Overrides: The full human edit flow, and the global active state management layer to pass those overrides to Cook Mode. - Single-Step Cook Mode Carousel: The high-contrast, distraction-free landscape step view featuring atomic inline text highlights for modified ingredients. Later: - Multi-Modal Ingestion Engines: The ability to import recipes via photo upload/OCR text extraction, local PDF processing, or native smartphone OS share-sheet extension inputs. - Ask Sammy to Substitute: engaging AI to suggest alternatives such as for buttermilk (milk + acid) that upon user selection, will insert new instructions to prepare the homemade buttermilk in the right order. - The Global Preference Learning Loop: Writing the backend database rules that track "Always Substitute" choices over time, transforming the app into a fully personalized, predictive kitchen profile. - Dynamic Smart Ingredient Scaling: Automatically recalculating entire multi-step recipe weights and volume metrics based on a user changing a single parameter (e.g., scaling a recipe down because they only have 2 eggs instead of 3). -Hands-Free Navigation UI: Voice-activated command protocols ("Next step", "Go back", "Show ingredients") to prevent users from touching the device screen with dirty kitchen hands. - Social & Community Infrastructure: The operational code for the final completion ritual screen—specifically, building native image galleries for the user's baked dishes, platform-wide recipe sharing links, and external social media distribution rails. | |||
| Initial Prompt Design | Master Prompt [Initial Design] | Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency? | Sammy Initial Master Prompt Updated 5/28/26 Refined Prompt Updated 6/7/2026 | |||
| Prepare for Testing & Iteration | Evaluation Criteria & Test Plan | Evaluation Criteria | What specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output? | Golden Evaluation Dataset 1. Zero Hallucination Baseline (Critical Safety) - Benchmark: 0% tolerance for hallucinated ingredients or numeric metrics. - Definition of "Good": The AI must never invent an ingredient, double an amount, or change a unit (e.g., transforming tsp into tbsp) that wasn't present in the source web scraping. 2. Tool Inference Accuracy (Relevance & Completeness) - Benchmark: ≥ 90% accuracy on extracting implicit kitchen gear. - Definition of "Good": When parsing the text instructions, the model must successfully log hidden tool requirements into the inferred_tools array based on context clues (e.g., "zest the lemon" must accurately output Zester/Microplane; "whip the egg whites" must output Whisk or Stand Mixer). It must not list irrelevant equipment. 3. Semantic Structure Validation (Clarity & Layout) - Benchmark: 100% adherence to the target component schema. - Definition of "Good": The AI must perfectly isolate multi-part recipes into logical visual components (e.g., separating "The Cake" raw string lines cleanly from "The Frosting" lines). Ingredients and steps cannot cross-contaminate across component cards. 4. Absolute Zero Fluff (Tone & Latency) - Benchmark: 100% Conversational Elimination. - Definition of "Good": The raw output must contain zero pleasantries, introductory remarks, or markdown prose commentary (e.g., no "Here is your parsed JSON payload!"). It must strictly output compile-ready JSON to minimize token delivery latency and prevent frontend component layout rendering crashes. 5. Numeric & Fraction Parsing Integrity (Data Accuracy) - Benchmark: 100% structural fidelity on mixed numbers and fractions. - Definition of "Good": The parsing layer must successfully format challenging recipe string fragments (like 1 3/4 cups or 1.75) as distinct, unbroken atomic units without dropping characters or truncating decimals, ensuring the text renderer highlights the values cleanly on the final Balsamiq wireframe layout cards. 6. Inline Instruction Refactoring Integrity (Precision Mapping) - Benchmark: 100% Error-Free Ingredient-to-Step Mapping. - Definition of "Good": When the AI rewrites raw instructions to insert ingredient metrics inline, it must achieve absolute precision across three distinct vectors: 1. Zero Numerical Drop: It must never omit an amount or truncate a fraction (e.g., 1 3/4 cups must never accidentally render as just 1 cup or 3/4 cup). 2. Contextual Accuracy: It must match the correct metric variant to the correct step execution block (e.g., if a recipe uses sugar in both the cake base and the frosting, the cake base step must show the cake base sugar amount, not the frosting sugar amount). 3. No Formatting Corruption: The injected values must consistently wrap cleanly inside our app's target string tokens so the Balsamiq wireframe carousel can apply the high-contrast text highlight pills without clipping the instruction prose. 7. Contextual "Chef's Tip" Extraction (Value-Add Capture) - Benchmark: ≥ 85% Recall Accuracy on high-value culinary advice. - Definition of "Good": The AI must successfully identify and extract non-procedural, critical baking wisdom buried in the blog text (e.g., "Make sure your eggs are strictly room temperature or the batter will break") and map it to a global chefs_tips array. It must never capture irrelevant personal blog fluff (e.g., "My grandmother used to make this on rainy Tuesdays"). 8: Contextual Density Scaling & Unit Translation - Benchmark: 98% accuracy against professional culinary weight charts. - Definition of "Good": Conversions must reflect volume-to-weight density realities (e.g., converting 120g of flour to 1 cup, but 200g of sugar to 1 cup) rather than flat scalar multiplication. Liquid metrics must round cleanly to standard kitchen utensils (e.g., 4.9ml simplifies to 1 tsp). | ||
| Example Cases | What specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs? | The AI Ingest Behavior: The prompt instructs the model to automatically add localized clarifying parentheticals for a US cook based on semantic context: Superfine Caster Sugar (200g / approx. 1 cup) and 3 Medium Eggs (UK Medium = US Large). The Perfect Case (Happy Path): A clean Epicurious URL with perfect JSON-LD metadata. 1. Happy Path Case: - Input: A clean, standard Epicurious baking URL containing valid Schema.org JSON-LD structural data. - Expected Output: Fast, deterministic extraction mapping title, source, ingredient arrays, and explicit steps perfectly into modular component cards with zero layout latency. 2. Complex Ambiguity Edge Case (The Chef's Logic Test): - Input: A raw text copy/paste block from a British baking blog using mixed metrics ("225g caster sugar", "3 UK medium eggs") and containing a hidden narrative tip ("Make sure your butter is room temperature or the emulsion will break!"). - Expected Output: The AI flags "ambiguity_detected: true", triggers the Smart Intake dual-measurement conversion view (mapping density accurately, e.g., 120g flour = 1 cup), extracts the implicit tools ("Whisk", "Sifter"), isolates the "Chef's Tip" cleanly from the blog fluff, and performs inline measurement refactoring into the steps. 3. Negative Failure Case: - Input: An expired/broken link, a non-recipe news article URL, or a text block completely missing ingredient quantity metrics. - Expected Output: The backend rejects the payload gracefully, completely preventing application processing loops or frontend screen freezes, and rendering a friendly kitchen-themed error state card to the user. | ||||
| DEVELOP | AI Model Selection & Justification | AI Model Selection & Justification | Which AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product? | SELECTED MODEL OpenAI GPT-4o-mini (Lean, High-Guardrail Architecture) CRITERIA & JUSTIFICATION 1. Multi-Layered Weakness Mitigation: While GPT-4o-mini has a lower native reasoning ceiling than flagship models, its limitations are entirely covered by our system architecture. First, Firecrawl executes the heavy data-cleaning layer, stripping out messy website HTML/ads and passing clean Markdown to the LLM. This drastically reduces token noise. Second, our Master Prompt enforces a strict Culinary Physics & Density Matrix for unit translation. Finally, the Screen 3 Human-in-the-Loop (HITL) Verification Canvas acts as the ultimate safety net, allowing users to catch and edit any minor parsing anomalies before saving. 2. Cost & Runway Efficiency: Utilizing GPT-4o-mini minimizes API operating expenditures by ~97% ($0.15 input / $0.60 output per 1M tokens) compared to GPT-4o. This provides a massive operational runway for rapid test iteration and evaluation loops during the capstone cycle. 3. Latency & Responsiveness: Mobile kitchen applications demand low cognitive waiting states. GPT-4o-mini processes unstructured payloads in under one second, ensuring an instantaneous, premium user experience during live deployment scoring. 4. Deterministic Output Structure: The model natively supports OpenAI's strict JSON mode, ensuring that the final data contract perfectly matches our frontend UI visual card structure with zero risk of conversational text leaks breaking the layout. | Samantha, the model selection defense is the strongest piece of Develop work here. Layering cost, latency, JSON mode support and architectural mitigation of each known weakness shows genuine trade-off reasoning and the eight numeric benchmarks give judges a coherent through-line across phases. The 85% tip isolation rate traced to a structural cause is honest test reporting. What undercuts all of it is sample size, three recipes tested against a harness built for fifteen to twenty-five means a single failure drops measured performance to 67% and claims like 98% density conversion accuracy cannot survive that math. Write ground-truth JSON for each of the five boundary conditions you named, label which criterion each case stresses and run the full set before Demo Day. For deploy start with your launch approach, the ten-user pilot already referenced in your eval plan, formalize it by naming who those users are and what a two week pilot looks like. Define success metrics as targets not names, something like 80% of pilot users successfully parsing three recipes in week one and fewer than 5% of parsed fields manually overridden. For AI monitoring track parse latency, override rate per criterion and API cost per parse. For legal document how you handle scraped recipe content since attribution policy matters when pulling from third-party blogs. Fill those fields with specifics and Deploy will carry its weight on Demo Day. | |
| Define Inputs | Input Specification Table | Required Fields | What are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement. | raw_input_source (String): The primary unstructured data string provided by the user on Screen 1. This must support two structural formats: - A standard web URL (which triggers the automated Firecrawl markdown scraping route). - A raw text copy/paste character block (the operational fallback bypassing third-party web firewalls). | ||
| Optional Fields | Are there any optional or user-customizable fields? How do they impact the AI’s output? | user_selected_locale (String [Enum: Original, US, Metric, Both]): A state-driven parameter passed to the data model ONLY if the engine returns 'ambiguity_detected: true'. This field is null at initial ingestion and is populated via the Screen 3 dialog box interaction. active_scale_multiplier (Float): A UI-driven scaling factor (Default: 1.0; e.g., 0.5, 2.0) managed entirely on the recipe card canvas. It dynamically triggers client-side instruction string recalculation before entering Cook Mode, without requiring a secondary ingestion loop. | ||||
| Define Good Output | Output Evaluation Checklist | Objective Criteria | What criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance) | To evaluate the extraction engine, outputs will be scored against a fixed 15–25 recipe golden evaluation dataset spanning five core boundary conditions (British-metric, Schema-less blogs, multi-component bakes, sterile/no-tip content, and expired URLs). 1. Deterministic Schema Adherence (Target: 100%): The output must strictly match the specified JSON data contract. Any missing arrays, broken object brackets, or leaked conversational text blocks that cause a frontend rendering failure constitute a critical fail. 2. Inline Instruction Refactoring Accuracy (Target: 100%): Measured against Eval #6. When an ingredient is localized (e.g., 200g caster sugar -> 1 Cup Superfine Granulated Sugar), Rule 4 must successfully inject the refactored measurement into the step text string. Zero instances of raw, un-localized source metrics are permitted in the instructions array post-conversion. 3. Core Ingredient Extraction Recall (Target: ≥98%): Every single critical ingredient required to physically execute the dish must be extracted into the JSON schema. Dropping a core ingredient or its base decimal value counts as a failure. 4. Inferred Tools Precision (Target: ≥90%): Semantic identification of required kitchen equipment based on instruction verbs (e.g., "sift" -> Sifter, "whisk" -> Whisk). Tools must have direct contextual utility; hallucinatory or generic tool additions drop this score. 5. Culinary Density Matrix Conversion Fidelity (Target: 98%): Volume-to-mass translations must utilize ingredient-specific density thresholds (e.g., 120g flour = 1 Cup) rather than a uniform mathematical multiplier, verified line-by-line against the golden dataset values. | ||
| Subjective Criteria | Are there any criteria that require human judgment or qualitative assessment? | Because automated string matching cannot evaluate semantic intent, human-in-the-loop (HITL) grading will evaluate the following qualitative dimensions across the evaluation dataset: 1. Semantic "Chef's Tip" Filtering Precision: Evaluated against Prompt Rule 7. A human reviewer will verify that technical utility and kitchen reassurances (e.g., "batter will look runny...") are successfully preserved, while subjective blogger anecdotes ("reminds me of my grandmother") are completely discarded. 2. Language Tone & Clarity Consistency: Extracted instruction steps must maintain an authoritative, clear, and professional "Executive Sous Chef" tone. The refactored steps must read naturally as actionable line items without sounding disjointed or mathematically robotic post-injection. 3. Extraneous Noise Leakage: Verification that no structural "web junk" (e.g., newsletter signup phrases, social share prompts, or copyright strings) accidentally bypassed Firecrawl and the LLM boundaries to land on the frontend recipe card layout. | ||||
| Prompt Design Iteration | Master Prompt [Final Design] | Prompt Version 1 | What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints. | The initial system prompt baseline for GPT-4o-mini has been isolated into a dedicated development document to maintain clean version control and rapid iteration tracking. Master Document Link: Sammy Master Prompt - Version 1.0 (Initial Dev Baseline) Techniques Utilized: - Strict JSON-mode enforcement for UI schema stability. - Mass-to-Volume Culinary Density Matrix boundaries. - Active Inline Instruction Refactoring injection rules (Rule 4). - Actionable Chef's Tip vs. Blogger Narrative Fluff filtering boundaries (Rule 7). | ||
| Prompt Iterations | If revised, what changes did you make and why? How do you track and record prompt evolution? | Structural Evolution: Refactored the core JSON schema from an abstract flat text string layout to a highly structured array model, nesting the tip property directly inside the individual instruction object indexes. This directly eliminated an early failure mode where the LLM would drop or clip implicit culinary execution warnings hidden within long prose body paragraphs. Constraint Fine-Tuning: Shifted from clinical text sanitization rules to an explicit verbatim voice retention model. This ensures the engine captures the original author's conversational brand tone for technical warnings while maintaining strict fraction text formats ("1 3/4", "1/3") to guarantee data field stability during active user editing loops. Versioning Controls: Maintained frozen prompt states within our technical tracking index (v1.0 through v1.2 Baseline). Prompt integrity is continuously verified via targeted human smoke tests against our core "chaos vector" test files before updates are committed to the live application scraper pipeline. | ||||
| Data Preparation & RAG Implementation | Data Preparation & RAG Implementation | What data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information? | 1. Dataset Scale & Scope: A dedicated 'Golden Evaluation Set' consisting of 15–25 highly diverse recipes will be curated manually. This set will deliberately span 5 target operational boundary conditions: (a) British-Metric terminology, (b) Schema-less/deeply nested blog prose, (c) Multi-component baking structures (e.g., separate crust vs. filling steps), (d) Sterile text with 0 tips, and (e) Expired/broken URLs. Each case will be explicitly labeled to stress-test specific performance benchmarks: Schema-less prose isolates our Step-Level Tip Extraction metric, while British-Metric and Multi-component bakes stress-test our Inline Quantity Injection and Density Conversion boundaries. The full set will be executed to mathematically validate our baseline 85% accuracy target prior to submission. 2. Data Preparation & Labeling: For each test case, the raw text/URL payload will be logged alongside a manually engineered 'Ground Truth' JSON target. This target will explicitly define the perfect expected output, including exact inline metric refactoring strings (Rule 4) and human-verified 'Chef's Tips' stripped of narrative fluff (Rule 7). 3. Performance Benchmarking: During the Development phase, prompt iterations will be executed against this 15–25 recipe harness. Success metrics will be tracked quantitatively across our 5 Objective Criteria benchmarks (Row 33) to prevent regression as the prompt is hardened. | |||
| Create Evaluation Set | Example Input/Output Data for Testing | Typical Examples | What are the most common inputs and expected outputs? Use real data if possible. | 1. Input Vector Example: Standard web URL from a mainstream food blog (e.g., Sally's Baking Addiction) passed directly from the user interface to the Firecrawl scraping layer. 2. Expected Pipeline Behavior: - Firecrawl successfully strips out HTML structures, navigation elements, and advertisement scripts, delivering a clean Markdown text block to GPT-4o-mini. - The model parses the data with zero ambiguity detected (ambiguity_detected: false). - The JSON data contract is cleanly populated with separated components, inferred tools, and instructions. - The UI canvas instantly renders the clean recipe cards on Screen 3 without launching the manual metric localization dialog wrapper. | ||
| Edge Cases & Negative Cases | What examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain) | 1. Edge Case: Deeply Nested / Schema-less Prose - Failure Mode: The recipe text is entirely unstructured and buried inside a massive personal blog post with zero metadata markers. - Mitigation: Prompt Rule 7 forces the model to ruthlessly drop narrative prose blocks, extracting only strings matching ingredient/instruction boundaries. 2. Edge Case: Aggressive Scraping Firewalls (e.g., NYT Cooking Paywall) - Failure Mode: Firecrawl returns an access-denied error or an empty markdown block due to cookie/paywall barriers. - Mitigation: The application interface detects the empty payload and automatically reveals the Screen 1 raw text copy/paste fallback box, allowing the user to bypass the web scraper entirely. 3. Edge Case: Absolute Measurement Chaos - Failure Mode: A recipe listing fluid ounces, grams, and vague descriptors ("a knob of butter", "juice of half a small lemon") in the same text string. - Mitigation: Prompt Rule 6 automatically triggers (ambiguity_detected: true), maintaining schema stability while launching the Screen 3 localized 4-way conversion dialog to let the cook declare their target baseline. 4. Implicit Technical Transitions: Sub-steps containing hidden "chemical execution guardrails" or sensory indicators without explict markdown delimiters or visual callout tags (e.g., "don't panic, it's supposed to look slighly curdled at this stage") | ||||
| Test Example Data & Review Results | Manual Review | Run your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why? | Manual Tests: Ran 3 recipes through the playground: "Peanut Butter Banana Cookies (mid-complexity-moderate narrative), "Love at First Sight Chocolate Cake (high complexity-high narrative), and "Turlu Turlu (low complexity-low narrative) Performance: Exceptional data extraction precision across core semantic mappings, including dynamic fraction string insulation and unitless variable count handling (e.g., parsing item numbers cleanly without leaking backend technical null strings). Failures &Observations: - The Cake case exposed a clear zero-shot cognitive bottleneck. Because the author embedded an execution warning ("the batter will be liquidy but that's a good thing") implicitly within an instruction sentence, the model failed to isolate it into the dedicated tip property, instead leaving it bound to the raw instruction string. - The "Chef's Tip" was sometimes isolated properly, but some narrative instructions were summarized (marking a potential optimization opportunity for hardening the prompt/logic, otherwise the intent was maintained). | |||
| Automated Evaluation | What pass/fail rate or scores did the AI achieve on core criteria? | Target Benchmarks: Post-baseline validation checks show a 100% success rate on preserving fraction string formats ("1 3/4", "1/3") for frontend form stability, and a 100% success rate on retaining the author's original verbatim phrasing for isolated tips. Accuracy Matrix: Our target benchmark is to achieve an 85% success rate on isolating implicit, paragraph-buried culinary warnings once executed across the full 15–25 recipe dataset. This establishing phase serves as our primary quality score before implementing few-shot data injections or automated structural validation scripts. | ||||
| Handle Edge Cases & Iterate | Edge Case Identification | What edge cases did you identify in testing or real usage? | 1. Implicit Text Anchors: Food blogs that wrap high-stakes chemical warnings directly inside narrative instructions without explicit visual markdown boundaries (e.g., "don't panic, it's supposed to look curdled here"). 2. Complex Unitless Quantities: Items that depend heavily on non-standard size descriptors (e.g., "juice of half a small lemon" or "6 thick artisanal sourdough slices") where a strict numerical integer is missing from the source text. 3. Multi-Component Scope Creep: Recipes that mix multiple standalone substructures (e.g., a pie crust, a cream filling, and an optional fruit topping) into a single unorganized text wall. 4. Mixed measurement units: Recipes that mix metric and US standard measurements. (The units render correctly, but appears to be a future UX enhancement opportunity) | |||
| Updates & Adjustments | What prompt or system adjustments have you made based on failures, feedback, or edge case observations? | Author's Voice Optimization: Hardened the behavioral guidelines to mandate that the tip property must capture text verbatim to retain original brand personality, rather than summarizing it clinically. Schema Realignment: Structural prompt updates were made to explicitly map the tip container inside the individual instruction step object loop. This guarantees 1:1 data contract alignment with our native front-end types (src/lib/recipe-types.ts), allowing Lovable to natively render the yellow callout boxes without custom regex string splitting. This adjustment successfully isolated the hidden paragraph warnings we flagged as edge cases in Row 42. | ||||
| Automate Evaluation Approach | Evaluation Method | What is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets? | HYBRID EVALUATION METHODOLOGY & PRODUCTION SCALING 1. Baseline Phase (Human-in-the-Loop): 100% manual grading executed across the 15–25 recipe golden dataset, scoring outputs line-by-line against our 5 Objective Criteria to establish ground-truth performance before go-live. 2. Pilot Phase Scaling (Telemetry-Driven): During the (planned) 10-user pilot, evaluation will scale via programmatic error logging. The Lovable frontend database will automatically capture instances where users manually edit or override parsed data on the Screen 3 Verification Canvas. 3. (Planned) Model-Graded Scaling (V2 Pipeline): For post-pilot scaling, a secondary 'LLM As A Judge' pattern will be introduced. A flagship model (GPT-4o) will be utilized to programmatically compare Sammy's runtime outputs against the ground-truth dataset, executing automated batch evaluations to eliminate manual human constraints. | |||
| Evaluation Frequency | How often will you re-run evaluations for new data, new prompts, or post-launch monitoring? | EVALUATION CADENCE & REGRESSION TESTING TIMELINE 1. Core Development Phase: Evaluations will be re-run manually after every major prompt iteration or rule adjustment to calculate performance drift and protect baseline accuracy thresholds. 2. 10-User Pilot Phase: A weekly evaluation review cycle will be established. User-corrected data payloads from the application's telemetry logs will be extracted every 7 days, formatted into new edge-case testing data, and used to expand the golden evaluation set. 3. Post-Launch & Production: Automated regression evaluation loops will execute as a mandatory continuous integration (CI) gate prior to any live production deployment of new prompt versions, upon migrating to a new underlying frontier foundation model), or backend model migrations to capture behavioral regressions.. | ||||
| DEPLOY | Finalize Launch & Rollout Plan | Operational Readiness Checklist | Technical Readiness | Is infra (APIs, databases, rate limits, monitoring, rollback) tested and documented? | Infrastructure Status: Yes. While the Sammy prototype is fully managed within Lovable: (diagram) - Frontend: React 19 (via TanStack Start) & TypeScript - Backend & Database: Supabase (Auth, Database, Storage) - Deployment/Runtime: Cloudflare Workers (Serverless/Edge) - Styling & UI: Tailwind CSS & shadcn/ui the target production infrastructure has been designed and documented (link). After the Pilot, I plan to transition Sammy from a the Lovable prototype to a self-hosted stack so I can control and manage the codebase, system prompt, and full telemetry pipelines, deployments and customizations. Technical specifications and data flow schemas will then be hosted in a centralized engineering Wiki and linked in GitHub. Core Components & Database: The ingestion pipeline utilizes Firecrawl for web scraping (now and future), routed into an Google Gemini 2.5 Flash (now) completion loop, with state data stored in a Supabase PostgreSQL instance (both now and future). In the future, we will transition to the OpenAI gpt-4o-mini model and precise System Prompt and Eval set. Rate Limiting: To prevent API token draining and control server costs, the production backend will enforce a hard limit of 5 recipe parses per user per hour, paired with a frontend button-debounce to block duplicate submissions. Timeout constraints: Ingestion tasks will enforce a hard 15 second backend timeout threshold to prevent hanging server processes. Limits are handled gracefully with proper user messaging and error handling if a source (URL) fails to respond. Multi-phase Monitoring & Rollback Plan: Prototype & Pilot (Lovable): Ingestion telemetry and failure rates are monitored programmatically via Supabase db logs that track user text overrides and schema parsing failures. This is a manual process for code updates and rollback procedures. Production (Future): Post migration, monitoring will integrate an LLM telemetry wrapper (e.g. LangSmith) to audit realtime LLM token usage, prompt latency, and processing bottlenecks. Rollback Strategy: In the future, the codebase will live in a managed GitHub repo. Deployment version rollbacks will be managed via GitHub Releases, allowing an immediate revert to the previous stable main branch commit if a production schema or prompt update breaks production stability. | Deploy section carries real operational weight, with explicit promotion criteria and the attribution-first legal posture that shows you have thought about the content-creator relationship before it becomes a problem. You are running Gemini 2.5 Flash now but plan to migrate to GPT-4o-mini, and there is no comparative eval bridging those two models. Every accuracy claim in Develop was generated against one model, and every Deploy metric assumes the other. Run your golden dataset against both before Demo Day so judges see a migration plan grounded in evidence, not intent. Second, tighten the pilot success criteria into weekly checkpoints so you know at day fourteen whether to proceed or iterate, rather than discovering it at day twenty-eight. |
| Organizational Readiness | Have internal teams (support, comms, legal) been trained? Is documentation complete? | Organizational Team Status: Since this is a solo venture, standard enterprise staff training is substituted with a centralized Operational Playbook and compliance policy documentation, which will be hosted in the GitHub project repo wiki prior to public launch. Comprehensive system documentation, ingestion rules, and compliance criteria will be hosted in the GitHub project repository wiki prior to public launch. This serves as the definitive onboarding and 'self-training' matrix for possible future external engineering or support contractors, minimizing manual operational overhead. Legal & IP Attribution Policy: To protect content creators and maintain strict copyright compliance when parsing third-party food blogs, the production application will enforce an 'Attribution-First' UI constraint. Every parsed card will permanently feature a prominent, un-editable hyperlink routing users back to the author's original source URL. The ingestion engine will deliberately omit narrative essays, ads, and original image caching—parsing structured text solely for personal cooking utility to maintain fair-use compliance. Please see Legal Policy Doc Automated Support & Triage Playbook: Customer support will be handled asynchronously via automated product telemetry. A 'Report Error' interaction on the User Verification Canvas will allow pilot users to instantly route parsing anomalies, schema mismatches, and raw source URLs into a dedicated Supabase incident_logs table. This replaces traditional support queues with a direct, data-driven engineering triage pipeline. | ||||
| Launch & Rollout Strategy | Launch Approach | What is your launch approach? Pilot, AB test, or all users—who gets access and when? | Prototype Phase 1 (The 'Now' Strategy): Approach: A 4-week, tightly controlled Closed Beta Pilot utilizing our live Lovable build to validate user behavior before executing a code migration. I will use a weekly checkpoint routine. To eliminate downstream deployment risks, the 10-user pilot is structured around weekly evaluative gates rather than waiting til the end of 4 weeks: - At the end of each week, I will perform the following checks and remediations: - (Ingestion Fidelity): Audit the first (up to 50) recipe ingests against Rule 5 (Numeric Constraints) and Rule 9 (Unit Collapse). Milestone: 0% decimal hallucination baseline on non-standard units. If failed, test the decoupled GPT-4o-mini System Prompt via the eval harness, and tune as needed for future deployment. - (Inference Utility): Conduct contextual interviews specifically tracking Rule 3 (Tool Inference Engine) accuracy. Benchmark: ≥90% user utility on auto-generated gear checklists. - (Ergonomic Adherence): Measure daily active usage (DAU/WAU/MAU) goals and track 'Back-to-Safari' user abandonment rates when executing the 'Step 0' global prep blocks. - (End of Pilot Evaluation): Compile qualitative feedback logs, freeze the pilot codebase, and finalize the technical specifications for the V1.1 decoupled production migration. Who: Restricted strictly to a single, high-touch cohort of 10 active users (Friends & Family) matching our primary target persona: Busy Home Cooks who routinely navigate online blogs, physical cookbooks, and handwritten cards. When & Feedback Loop: Launching immediately following internal verification testing. Users will receive personal email invitations with instructions to submit feedback directly via native in-app logs or direct triage channels. Strategic Intent: This 4-week window allows home cooks a realistic meal-planning and grocery cycle to test Sammy, giving us critical baseline data on ingestion friction, manual text overrides, and mobile UI performance. Production Phase 2 (The 'Future' Strategy - When Sammy Becomes a Real Boy): Cohort Overlap: Upon migrating to our independent full-stack architecture, we will execute a slow, deliberately managed production pilot. We will re-onboard select users from the original prototype cohort to act as a benchmark group, allowing them to identify performance gains, celebrate enhancements, and establish a highly engaged community foundation. The 'DM for Access' Rollout: The volume gate will expand via a curated LinkedIn community release utilizing a selective "DM me for access" strategy to tightly manage feedback loops and system load. The 10x Scaling Gate: Once platform telemetry stabilizes at 100 active users with net-positive accuracy metrics, Sammy will execute a 10x scale target to 1,000 users by unlocking organic social network sharing mechanics. | |||
| Scale Readiness | How will you ensure readiness for scale? How will you monitor initial volume and scale up? | 1. Scale Verification Strategy: Scale readiness will be managed through a phased, gate-controlled volume expansion. This ensures we thoroughly validate application logic and server stability before committing capital to full-stack independent infrastructure. 2. Initial Volume Monitoring (The Prototype Phase): - The Metric: Volume is strictly capped at our 10-user pilot pool for 4 weeks. - The Telemetry: Because we are inside the managed Lovable framework, we will monitor volume and stability programmatically via our Supabase database tables. We will actively track: - Total API token consumption rates per user. - The frequency and specific fields of manual text overrides on the User Verification Canvas (to isolate prompt degradation or scraping failures). 3. Scale-Up Operational Gates (The Production Phase): - Gate 1 (To 100 Users): Progression to our LinkedIn select release requires a baseline of 0 server crashes during the pilot, combined with an automated ingestion accuracy rate of 95% or higher (meaning minimal user adjustments needed on the canvas). - Gate 2 (To 1,000 Users & Beyond): Prior to unlocking viral social sharing, Sammy will exit the Lovable sandbox to deploy our independent React/Next.js production architecture. At this stage, manual database audits are substituted with advanced cloud monitoring, integrating LangSmith for real-time prompt/model auditing and specialized Supabase telemetry to track database connection limits and latency bottlenecks under load. | ||||
| Go-to-Market Plan | Marketing / Training Assets | What assets (FAQ, demo, guides) will you prepare for external communication/marketing? | 1. Prototype Phase Assets (The 10-User Pilot): The Personalized Onboarding & Quickstart Email: A high-touch, lightweight onboarding email sent directly to our 10 testers. This email will deliver the live Lovable mobile application link alongside a lightweight, bulleted 'Quickstart Guide' detailing exactly how to navigate the User Verification Canvas and how to report anomalies directly via the native 'Report Error' button, and specific testing boundaries (e.g., 'Please test at least one web URL and one physical photo upload this week'). 2. Production Phase Assets (The 100 to 1,000 User Scale-Up): The Interactive 'Feature Tour' Asset: A lightweight, automated frontend onboarding carousel (using a tool like Intro.js) to walk new public users through the dual-ingestion pipeline flow seamlessly. The Collaborative Product Demo Video: A crisp, 60-second screen-recorded walkthrough (Loom format) hosted on our landing page, demonstrating the speed of moving from a messy food blog URL to a clean cooking canvas. The 'Attribution-First' Creator FAQ: A public-facing documentation page explicitly detailing Sammy's fair-use compliance posture, data deletion procedures, and direct contact forms for publishers, ensuring absolute transparency as community adoption scales. | |||
| Stakeholder / Internal Comms | How will you communicate launch plans, progress, and outcomes internally? | 1. Stakeholder Definition: As a single-founder venture, traditional cross-departmental corporate communications are substituted with a streamlined reporting loop managed exclusively for our advisory circle, strategic mentors, and key design partners. 2. Communication & Reporting Mechanics: > * Pilot Progress Reporting: Launch milestones and pilot performance data will be synthesized asynchronously (roughly) every 14 days into a lightweight, bulleted 'Founder's Update' and posted to internal documentatioin repos. Data Synthesis: This update will distill core metrics directly from our Supabase database tables (specifically total recipes parsed, system latency, and user text-override frequencies) to transparently highlight technical progress and prompt engineering stability. 3. Feedback Integration Loop: Insights, feature requests, and strategic feedback gathered from advisory sessions will be centralized in our master product backlog, ensuring external expert guidance directly shapes our post-pilot independent production migration roadmap. | ||||
| Confirm Legal, Privacy & Risk Protocols | Data & Privacy | How do you handle and protect user data, including storage, privacy, and compliance? | 1. Data Storage & Segregation: Data handling is strictly bifurcated based on the ingestion vector to optimize privacy and security controls: - Public Ingestion (URLs): Web-scraped content is processed textually and stored as open-schema JSON blocks within our primary database. No local copies of creator media assets are retained. - Private Ingestion (User Photos): User-uploaded photos of physical cookbooks or handwritten cards are treated as highly sensitive, private assets. These files are securely hosted inside isolated Supabase Storage Buckets protected by strict Row-Level Security (RLS) policies. 2. Privacy Compliance & Access Control: > * User Isolation: Supabase RLS protocols ensure that user-uploaded media files are programmatically restricted exclusively to the authenticated user ID that executed the upload. Third-party indexing, public search caching, and cross-user data exposure are strictly blocked by the database contract. - Zero Training Footprint: In alignment with modern privacy expectations, user-uploaded proprietary text layouts and custom recipe images are explicitly excluded from public LLM fine-tuning or training datasets. 3. Data Retention & Deletion: Users retain absolute data sovereignty. Deleting a generated recipe card triggers a cascading backend delete command, permanently wiping the structured schema from the PostgreSQL instance and erasing the original media file from the Supabase bucket. | |||
| Policy & Compliance | Are content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain? | Yes. Programmatic compliance and legal guardrails are active for the prototype phase and fully mapped for future production scale. Current Prototype Safeguards (The 'Now' Strategy): - Legal Compliance: Ingestion logic strictly adheres to our established Intellectual Property & Fair Use Policy Draft—separating un-copyrightable factual data from protected creative expressions. - Content Moderation: Input risk is low as Sammy functions as a personal utility. Firecrawl programmatically sanitizes incoming website HTML to strip out malicious scripts before text hits the LLM layer, and OpenAI's native safety API filters block non-culinary text processing. - Audit Process: System accuracy is audited manually through direct inspection of the Supabase data tables to evaluate text-parsing stability and track any user-reported canvas errors. Production Safeguards (The 'Future' Strategy): - Continuous Ingestion Auditing: Upon full-stack migration, manual database audits will be replaced with LangSmith telemetry pipelines to automatically track prompt drift, hallucination rates, and LLM output accuracy at scale. | ||||
| Define Success Metrics | Success Metrics | User/Business Metrics | What user metrics will indicate success? What business metrics will demonstrate value? | User Success Metrics (Engagement & Utility): - Ingestion Completion Rate: The percentage of initiated URL scrapes or photo uploads that successfully result in a saved recipe card (Target: >90%). Monthly Active Cooking Sessions (Retention): Percentage of users who open at least 1 saved recipe card in "Cook Mode" per month, proving the app is a reliable kitchen companion when they do decide to cook (Target: >60% of pilot cohort active monthly). Business Value Metrics (Viability & Scale Efficiency): - Average Token Cost per Parse: Tracking the exact financial cost of OpenAI API calls per recipe ingestion to ensure the unit economics scale safely before launching a premium tier (Target: < $0.02 per parse using gpt-4o-mini). - Viral Coefficient (K-Factor): The average number of new users generated organically by a single active user sharing recipe cards. While true sustainability scales at K > 1.0, our Phase 2 pilot success milestone is set at K > 0.5 (meaning every 10 users organically bring in 5 new cooks), significantly reducing our customer acquisition costs before opening public social channels. | ||
| AI Metrics | How will you measure AI performance and accuracy? | AI Performance Evaluation Strategy: AI accuracy will be graded using a combination of behavioral user telemetry (the 'Now' phase) and automated programmatic assertions (the 'Future' phase) to isolate model degradation and optimize prompt stability. Prototype Phase Metrics (The 'Now' Strategy via Supabase): - Canvas Correction Rate (Friction Metric): Tracking the percentage of parsed ingredient rows or steps that a user manually edits on the Verification Canvas. A high correction rate indicates prompt drift or an unoptimized model layout (Target: < 10% of fields corrected per recipe). - Ephemeral Image Ingestion Success Rate: The percentage of user cookbook photo uploads that successfully map to a complete JSON recipe schema without triggering a fatal LLM parsing error. Production Phase Metrics (The 'Future' Strategy via LangSmith): - Hallucination & Omission Auditing: Utilizing LangSmith evaluator assertions to compare raw markdown web scrapes against the generated JSON schema to programmatically flag if the model invented an ingredient or omitted a critical cooking step. - Latency Threshold Compliance: Measuring the end-to-end LLM processing speed (time-to-first-token and total completion time) to ensure the entire AI pipeline remains strictly under our hard 15-second backend timeout threshold. | ||||
| Monitor, Iterate & Improve | User Support & Feedback Plan | Support Channels | Where can users get support? Is escalation and ownership clear? | Prototype Phase (The 'Now' Strategy): Support channels are entirely bypassed to minimize development overhead. For our 10-user pilot, support is executed via a scrappy, high-touch email and in-person channels. Testers (friends & family) will submit bugs, feedback, and scraping failures directly via email or in-person. This allows for manual triage and direct, qualitative user interviews during the critical initial validation cycle. Production Phase (The 'Future' Strategy): As Sammy scales from 10 to 100 to 1,000 users, support will be added to transition to a three layer system to completely eliminate manual inbox triage: 1. In-App 'Report Issue' routing: A native button within the UI allows users to instantly report a problem, programmatically writing the error metadeta directly to an 'isolated incident_logs' table in the Supabase instance. 2. Support chat (Crisp.chat): A lightweight, embedded in-app chat widget used exclusively for automated deflection and async context gathering. A chatbot workflow will handle immediate, common troubleshooting and FAQs. If a human is needed, the widget will act as an inbox, communicating 24hr response time expectations and automatically capturing the user's device logs and other data (URLs) for review later. 3. Automated Service Desk (Linear): High severity bugs and parsing anomalies that cannot be resolved via chat will will drop down to the Linear backlog via webhook, This instantly creates a dev ticket linked directly in the GitHub repo for planning. | ||
| Feedback Workflow | How do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated? | Pilot Phase (The 'Now' Strategy): Feedback Gathering & Prioritization Framework: During the PIlot, feedback is gathered from pilot users directly via email or in-person, and final qualitative interviews after a cook session. Incoming bugs are triaged purely by severity and impact on the core experience. (see prioritization framework below). Release Management & Communication: Critical fixes are deployed directly in the prototype, and resolution is communicated immediately to the affected user via email to confirm resolution and invite them to re-verify the recipe in their kitchen. Production Phase (The 'Future' Strategy), As Sammy scales, manual gathering transitions to a passive, automated model: quantitative data is collected via automated database logs when an ingestion fails, while qualitative bugs drop into a Linear dashboard via in-app widgets. Release Management & Communication: As Sammy scales, manual communication will be replaced by a transparent "Build in Public" digital hub: - App Store/Play Store Changelogs: Standard technical release logs will accompany all major production builds. - Sammy Public Roadmap & Release Webpage: A centralized, user-facing webpage will host a simplified, dynamic timeline of "Latest Releases" and "Upcoming Fixes." This allows cooks to check for app updates independently, reducing duplicate bug submissions. Prioritization Framework: - Critical/High: System crashes, scraper time-outs, or fatal schema generation issues that completely block a user from saving or executing a recipe. - Medium (Prompt Drift): Specific recipe domains or blog layouts that consistently force users to manually edit fields on the Verification Canvas (triggers prompt optimization). - Low: Aesthetic layout adjustments, UI polish, or cosmetic feature requests. | ||||
| Monitoring & Continuous Improvement | Monitoring Approach | What monitoring/logging is in place to spot operational/AI issues post-launch? | Prototype Phase (The 'Now' Strategy): Monitoring is kept lightweight and centralized inside our backend. We utilize basic Supabase Database Logs to track live system health. Every time a user triggers an ingestion, the system logs the outcome. This allows the founder to directly monitor the database to catch fatal server crashes, scraping timeouts, or API connection errors. Production Phase (The 'Future' Strategy): As traffic scales, manual database checking will be replaced with automated observability pipelines: Application Monitoring (Sentry): To automatically capture and alert the founder to front-end UI crashes or silent Javascript exceptions in the kitchen. AI Quality Telemetry (LangSmith): To automatically monitor LLM token spend, latency spikes, and prompt accuracy drift across thousands of concurrent recipe parses. | |||
| Ongoing Improvement | How will you collect learnings, review performance, and update your system continuously post-launch? | Prototype Phase (The 'Now' Strategy): System learnings are derived directly from the Verification Canvas override logs. The founder reviews the specific fields users are forced to manually correct, pinpointing exactly where the parsing prompt failed. Prompts are updated directly in the code stack, verified locally, and pushed to the live prototype. Production Phase (The 'Future' Strategy): Continuous improvement scales into a structured Git-Controlled Workflow: - Prompt Versioning: Prompt iterations will be decoupled from core application code, managed and regression-tested inside LangSmith to ensure new tweaks don't break existing parsing logic. - Isolated Deployments: All system updates, database schema changes, and UI fixes will be built on isolated GitHub feature branches and run through automated testing environments before being deployed to live users, ensuring zero downtime for cooks in the kitchen. | ||||




