Consumer Parenting
Pebble
Pebble is an AI-powered recommendation app for parents trying to find child-friendly places through trusted social connections instead of scattered listings and ads. Parents upload a photo and a few notes, and the system infers the location, drafts a recommendation, and shares it back into a searchable friend-based network. The product starts with toddler parents in Luxembourg and expands toward broader family decision categories such as schools, camps, and restaurants.
The problem
Parents deciding where to take a toddler face a scattered, low-trust search. Gathering options across Facebook groups, WhatsApp threads, event calendars, blogs, and Google Maps takes 10–30 minutes and returns a long, redundant list of generic places. Worse, the signal that matters is missing. Smaller local venues have little or no rating, and where ratings exist, "4.5 stars from 200 strangers" says nothing about whether a place works for one specific 19-month-old. Good recommendations from trusted parents live in WhatsApp and drop-off conversations, and evaporate between the Monday they're mentioned and the Friday they're needed.
The solution
Pebble is an AI recommendation app that turns a parent's visit into a structured, shareable record and answers questions with trusted, friend-based results. A parent uploads a photo and a few keywords, and Pebble drafts a recommendation the parent confirms in their own voice. On retrieval, natural-language queries return 1–3 recommendations tagged by who they came from — a friend, a same-crèche parent, a parent of a same-age child — rather than 50 generic results. Every logged place is an endorsement by definition, so there are no star ratings to interpret; the parent decides who to trust based on the relationship. Pebble starts with toddler parents in Luxembourg and the Greater Region.
How it works
V1 runs two models server-side. Claude Sonnet 4.6 handles contribution: a forced `submit_recommendation` tool call reads the photo and any contributor text and returns validated JSON, with place identity trusted only from Google Places — vision is treated as a signal, never the source of truth, to prevent fabricated details. Two small compose calls turn the parent's answer into a first-person note and an optional heads-up tip. Semantic search uses Voyage's voyage-3.5 embeddings at 1024 dimensions, matched through a Supabase `match_places` function whose `SECURITY INVOKER` setting keeps the friends-only access policy active. The trust-critical retrieval path is deterministic, so a model outage degrades polish and relevance but never trust.
Who it's for
Pebble is B2C, built for the "Tired Friday Parent" — a parent of a child aged 12–48 months in Luxembourg or the Greater Region. This persona skews foreign-born (matching Luxembourg's ~47% average), arrives with few local connections, captures 20+ photos per outing, and is comfortable with AI-chat interfaces. Planned expansion reaches parents of preschool, school-age, and tween children. A secondary B2B2C path envisions crèche partnerships and, later, verified venue listings — provided verification never becomes promotion and undermines the trust mechanic.
Why it matters
Friend-recommendation apps are historically a graveyard — Wist, Path, and Foursquare Tips all failed on the cold-start problem and the free, zero-friction substitute of the camera roll and group chats. What has changed is that AI collapses contribution cost: a photo plus two keywords now produces a useful structured entry, reopening the category. Luxembourg's dense, high-income, weakly-networked expat parent population is uniquely suited to seed it, and the personal-log-plus-friend-graph pattern is already validated by Letterboxd, Beli, and Goodreads. Pebble is at idea-validation stage, launching as a closed, invite-only pilot to 5–8 friend users, with expansion gated on trust-loop signal rather than a date.
The workflow
The PRD
| PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0 | ||||||
|---|---|---|---|---|---|---|
| Your Name: | James Cai | |||||
| Your Product: | Pebble | |||||
| Your Industry: | Digital parenting and family-lifestyle apps | |||||
| Date: | May 15, 2026 | |||||
| 4D Method | AI PRD | Instructor Feedback | ||||
| Phase | Activity | Theme | Topic | Key Question(s) | Your Response Include external links to visuals/prototypes as required. | |
| DISCOVERY | Understand your market, business, product & user context | Business Value Map | Market Attractiveness | What industry is your business in? (ie Financial services, Healthcare, Education, etc)? | Pebble sits at the intersection of three categories: digital parenting and family-lifestyle apps (the persona and core use case), consumer social discovery (the friend-graph trust mechanic, similar to Letterboxd or Beli), and local recommendation and community tech (the geographic concentration and crèche-based bootstrap). | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors? | **Headwinds:** - *"Friends recommend things" is historically a category graveyard.* Wist, Path, Foursquare Tips, Polyvore — every general-purpose friend-recommendation app of the last decade failed. The cold-start problem is severe: until enough friends are on the platform, the product has no value. - *The substitute is free and zero-friction.* For the memory function, the camera roll is automatic and already organized by date/location. For the discovery function, WhatsApp parent groups are one tap away. Any new product must beat both, not just one. - *Trust-based products are hard to defend.* Without genuine network density or proprietary data, "trustworthy recommendations" can be copied by larger players (parenting magazines, FB communities) with greater distribution. - *City-by-city local-network seeding.* Unlike pure-SaaS products, Pebble does not enter new geographies simply by translating marketing copy. Trust networks are inherently local — each new city requires seeding into local parent affinity structures, and the specific bootstrap unit that works (whether crèches, schools, workplace parent groups, neighborhood communities, or friend-of-friend chains) is itself a hypothesis to be tested per market. This constrains expansion velocity — though it is also part of the moat, since any competitor entering a new market faces the same friction. **Tailwinds:** - *AI collapses the contribution friction problem.* The historical reason these products failed — high cost of writing reviews vs. low retrieval value — has materially changed since 2024. A photo plus two keywords now produces a structured, useful entry. This is genuinely new and reopens a category. - *Luxembourg's parent population is uniquely suited.* Dense expat community with weak local networks (most expats arrive without family or established friends), high disposable income (Luxembourg GDP per capita ~€130K), and a trilingual+ culture creating multiple distinct affinity circles within a small geography. - *Recommendation fatigue and trust collapse in generic platforms.* Consumers are increasingly skeptical of Yelp, TripAdvisor, Google reviews, and similar generic recommendation sources — they are seen as gamed by businesses, polluted by AI-generated content, and stripped of the personalization that makes a recommendation actually useful. Parents experience this acutely: "4.5 stars from 200 strangers" tells them nothing about whether a place actually works for *their* toddler with *their* nap schedule. Demand for trusted, affinity-relevant recommendations is rising sharply across consumer categories, and Pebble's friend + same-kid-age mechanic addresses this directly. - *The model has been validated in adjacent verticals.* Letterboxd, Beli, and Goodreads have all built $50M–$500M businesses on the personal-log-plus-friend-graph pattern. Pebble applies a proven flywheel to a new vertical, reducing invention risk. | |||||
| What is the projected growth rate of your target market segment over the next 3-5 years? | Global family/parenting app category: ~15–20% CAGR (estimated $5B in 2023, projected ~$9B by 2027) - AI-enabled consumer apps subset: 30%+ CAGR through 2027 | |||||
| Business Model | What growth stage is your business currently in (e.g., startup, scale-up, mature)? | Idea-validation stage | ||||
| How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?) | **Primary model: B2C freemium subscription.** - **Free tier:** Core logging, AI-assisted contribution, friend graph access, basic AI semantic search, the ability to see your own pebbles and your network's pebbles within the same crèche/affinity radius. - **Premium tier (target €4–6/month or €40/year):** Unlimited entries, advanced semantic search across the entire network, annual printed memory book of your child's adventures (the "Spotify Wrapped for toddler parents" emotional hook), ad-free experience, expanded affinity radius (quartier + Greater Region), priority support. **Secondary models to evaluate post-validation:** - **B2B2C crèche partnerships.** Crèches and maisons relais pay a modest annual fee (~€500–1,000/year) for a branded parent-network instance of Pebble used as a community tool. This is both a distribution moat and a revenue stream. - **Verified venue partnerships.** Local toddler-friendly businesses (indoor playgrounds, family cafés, farms) pay to verify their listing's operational accuracy. Critical caveat: verification must not become promotion, or the trust mechanic collapses. To revisit only once the core friend-recommendation trust is firmly established. | |||||
| Who is your primary customer base (B2B, B2C, B2B2C)? | **Primary: B2C** — parents of children. v1 focus is parents of toddlers (0–4 years) in Luxembourg and the Greater Region because the discovery and trust gap is most acute in this band. Demographic expansion to preschool, school-age, and tween parents is the planned second axis of growth alongside geographic expansion. | |||||
| Differentiators | What are the key differentiators for your company? | 1. **AI low-friction contribution with preserved voice.** Photo + 2 keywords → structured entry (location, age-suitability, operational details) plus one human-voiced sentence. Collapses contribution cost while keeping the trust texture prior friends-recommend apps lost. 2. **Layered transparent affinity trust.** Recommendations tagged by provenance (friend / same crèche / quartier / same kid age). Tunable per query — not a single social graph. 3. **Personal-record first (Letterboxd/Beli/Goodreads pattern).** Log is valuable to the parent on day one without any friends present. Inverts the typical cold-start problem. 4. **Local-affinity organizational-unit bootstrap (current hypothesis: crèches).** Distribution via pre-bonded parent communities, not mass acquisition. Alternative units if crèche underperforms: schools, workplace parent groups, neighborhood communities, expat affinity circuits, friend-of-friend chains. 5. **Hyper-local Greater Region focus.** 90-min radius → 300–500 family destinations across Lux + Trier + Metz + Saarbrücken + Arlon. 6. **Pull-not-push UX.** No social feed. Opens at moments of need (Friday night, special occasion, friend asks). A tool, not a habit. | ||||
| Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.) | Customers | Who are the customers (ie buyers) of your product? | Skipped this question as this is a 0-1 product. | |||
| End Users | Who are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context? | Skipped this question as this is a 0-1 product. | ||||
| Current Products / Services | If you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers? | Skipped this question as this is a 0-1 product. | ||||
| User Value Map | Target Persona | Who is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc) | **Primary persona: "Tired Friday Parent."** - Parent of a child aged 12-48 months in Luxembourg or the Greater Region. - 47% probability the parent is foreign-born, matching the Luxembourg national average. Foreign-born parents typically arrive with fewer than 5 local family connections. - Household income exceeds €80,000 annually. Dual-income or single-income-high-earner. - Captures 20+ photos per family outing. - Maintains 2-8 local parent contacts. Primary contact sources: crèche drop-off, neighborhood, expat community. - Smartphone-native. Comfortable with AI-chat interfaces. **Secondary personas (planned, not version 1):** - Parents of preschool and school-age children (4-10 years). - Parents of tweens (10-14 years). - Multi-kid families spanning age bands. | |||
| Journey Map (current-state) | What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product? | 1. **Start thinking about the weekend.** Thursday or Friday evening, the parent realizes the weekend is two days away and starts thinking about where to take the toddler. 2. **Gather information — 10-30 minutes.** The parent opens 2-5 sources: Facebook parent groups, WhatsApp parent groups, vdl.lu calendar, plurio.lu, blogs, Google Maps. Scrolling produces a long list of generic options. 3. **Assess the information — 5-15 minutes.** The parent looks at each candidate place and tries to judge whether it actually fits their child this weekend. They check ratings, ask in WhatsApp, and try to weigh weather, nap schedule, age suitability, partner energy, and drive time. 4. **Decide where to go.** Most weekends, the parent defaults to a familiar place (the usual park, the zoo, the grandparent's house). New activities are the exception. (Hypothesis under test in parent interviews.) 5. **Go to the place.** Saturday morning, the parent drives out, spends the day, takes 20+ photos. 6. **Try to remember the place weeks later.** Within 4-6 weeks, the photo sits in the camera roll mixed with everything else. The name and details fade. 7. **Try to share the place with a friend.** The parent mentions the place in passing at drop-off or in a group chat. The friend hears it on a Monday, needs it on a Friday, and cannot recall the name. Both sides lose the information. | ||||
| Pain-points | Where does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe? | Friction concentrates at four steps of the seven-step journey. Steps 1, 4, and 5 are low-friction transitions. Steps 2, 3, 6, and 7 contain the unmet needs. **Step 2 — Gather information (highest friction).** Just gathering enough information to consider an option takes 10-30 minutes. Two pain points sit here: - **Information overload.** Parents open 2-5 sources (Facebook groups, WhatsApp groups, vdl.lu, plurio.lu, blogs, Google Maps) and get back a long, redundant list of generic options. Reason: each source publishes its own slice. Nothing aggregates them. The parent does the de-duplication and triage manually. - **Operational details scattered or missing.** Even after a candidate place is identified, parents must dig across multiple pages to confirm opening hours, parking, stroller access, changing tables, dog policy, and toddler-specific hours. Often the information is not published at all. Reason: listings and editorial sites describe what a place *is* (story, photos, address), not the operational data parents need before driving 45 minutes with a toddler in the back seat. **Step 3 — Assess the information (highest friction).** Once information is gathered, parents must judge which option actually fits their child this weekend. Two pain points sit here, both signal problems: - **Insufficient signal.** For many candidate places — smaller local farms, recently opened venues, niche destinations — there is little or no rating, review, or first-hand account to draw on. Reason: a long tail of family-friendly places in the Greater Region sits below the threshold that attracts large review counts. The parent is guessing. - **Untrusted signal.** Where ratings do exist, the parent cannot tell whether they apply to their child. Reason: 4.5 stars from 200 strangers averages across all visitor types. Stranger ratings strip the variables (child age, mobility, attention span, weather tolerance) that determine fit for one specific 19-month-old. Trusted sources — a friend with a same-age child, a parent from the same crèche — exist but live in scattered WhatsApp threads and in-person drop-off conversations, not in any queryable surface. The combined effect is decision fatigue and defaulting to the familiar park. **Step 6 — Try to remember the place weeks later (high friction).** - **No way to retrieve a past visit.** A few weeks after a good outing, the parent cannot recall the name or find the place again. Reason: the camera roll keeps the photo but not the structure. Photos are timestamped and geo-tagged, but not tagged by activity type, age suitability, or "would return." The only retrieval surface is scrolling backwards through time. **Step 7 — Try to share the place with a friend (high friction).** Casual recommendations do not survive the gap between mention and need. Two pain points sit here: - **Recommendations received slip away.** A neighbor mentions a great farm at drop-off. Two weeks later, the parent cannot recall the name. Reason: the recommendation arrives at a moment of low intent (Monday morning) and is needed at a moment of high intent (Friday evening). No capture surface bridges the two. - **Recommendations given do not land.** Parents who want to pass on a good place have no easy way to do so beyond ad-hoc messaging. Reason: the cost to share (write a paragraph, find the right WhatsApp thread, attach a photo) exceeds the perceived benefit. The recommendation rarely reaches the right friend at the right time. | ||||
| AI Opportunities | From your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first. | AI-solvable pain points, ranked: 1. **Not enough information or ratings to decide, and ratings that don't apply (Step 3).** AI pulls together recommendations from people similar to the parent — their friends, parents from the same crèche, parents in the same neighborhood, and parents with same-age children. These voices fill in places strangers never rated. Each recommendation shows *who* it came from, so the parent can judge for themselves whether to trust it. Output: 1-3 recommendations from people the parent has a reason to listen to, instead of 50 generic results from strangers. 2. **Operational details scattered or missing (Step 2) and no way to retrieve a past visit (Step 6).** A parent who has been to a place creates a recommendation from a simple photo and a few keywords. AI extracts structured information from that input — name, location, plus operational details such as opening hours, parking, stroller access, changing tables, and dog policy. That structured record becomes available for queries by the parent and by the network from then on. The next parent who asks gets the answer in one shot instead of opening five tabs. 3. **Information overload (Step 2).** AI conversational interface accepts natural-language queries (example: "place with animals near a forest where my 2-year-old can roam"). The interface compresses gathering and filtering into one step and returns a small curated set instead of a long list. 4. **Recommendations received slip away (Step 7).** AI structures and stores network recommendations as queryable knowledge. A friend's casual mention becomes a permanent retrievable record. 5. **Recommendations given do not land (Step 7).** AI auto-generates a draft entry from a photo plus 2 keywords. User reviews and confirms. Contribution time drops from 3-5 minutes to 15-20 seconds. **Pain points NOT well-served by AI:** - **Network distribution.** Getting enough friends on the platform is a go-to-market (GTM) problem, not an AI problem. - **Affinity density.** Per-geography user count requires real acquisition, not synthesis. | ||||
| Develop an AI Solution Hypothesis | AI Solution Hypothesis | Diverge | Ideate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage. | Twelve candidate AI solutions, grouped by which side of the journey they address. **Contribution side — turn a visit into a structured record:** 1. **Photo + a few keywords or a short sentence → AI-extracted entry.** Parent uploads a photo and types a few keywords or a short sentence ("Robbesscheier with goats and donkeys, great for under-3s"). AI extracts name, location, age-suitability tags, opening hours, parking, stroller access, changing tables, and dog policy into one structured record. 2. **Voice input with auto-transcription (both sides).** Parent talks to the app instead of typing — like ChatGPT voice mode. On the contribution side: a 10-second voice note on the drive home ("we went to Robbesscheier, my daughter loved the goats") becomes a structured entry. On the retrieval side: a parent with a toddler in one arm asks out loud "any place with goats near a forest?" and gets the same answer typing would have produced. Fits the hands-full reality of parenting toddlers. 3. **WhatsApp forward → AI-extracted entry.** Parent forwards a WhatsApp message ("you should try this farm!") to a Pebble bot. AI extracts a place from the message and adds it to the parent's log as a "received recommendation." 4. **Geolocation-triggered prompt.** Phone detects parent has been at a new location for 2+ hours. End of day, Pebble asks "Was this Hähnchenhof? Want to log it?" 5. **Calendar-aware prompt.** Pebble notices the parent had a free Saturday and visited somewhere new. Sunday morning, it asks for a 30-second log. **Retrieval side — turn a question into a short, trusted list:** 6. **Natural-language conversational query, single-turn or multi-turn.** Parent asks "place with animals near a forest where my 2-year-old can roam" and gets 1-3 matches from friends, same-crèche parents, same-quartier parents, and parents of same-age children. Parents can also have a multi-turn conversation about a place — follow-ups like "is parking easy?", "who else went there?", "is it OK for a one-year-old too?", "did Sarah mention anything about lunch options?" — and the AI answers from the structured records, the human-voiced sentences, and the friend layer. Conversation feels less like a search box, more like asking a knowledgeable friend. 7. **Trust-tagged result cards.** Each result shows who recommended it — for example, "Sarah, your friend, recommended this place. She went with her 22-month-old." Every entry on Pebble is a recommendation by definition: the act of logging a place is the endorsement. There are no star ratings to interpret. The parent decides who to trust based on the relationship, not on a numerical score. **Network and memory side — surface the network's existing knowledge:** 8. **Cross-friend gap-filler.** When a place has no friend signal yet, AI surfaces aggregated same-age and same-quartier parents' signal. 9. **Annual memory book.** AI generates a retrospective at year-end from the parent's entries — a Spotify-Wrapped equivalent for family weekends. 10. **Multi-language contribution and retrieval.** Each parent contributes in their own language (English, French, German, Luxembourgish, with Portuguese under consideration). When another parent searches in a different language, AI translates the recommendation, the human-voiced sentence, and the operational details into the asker's language. The original-language text is preserved alongside the translation. A French-speaking parent's farm recommendation reaches an English-speaking parent without either of them switching languages. | ||
| Converge | Rank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project. | | Rank | Solution | Impact | Feasibility | Net | | ---- | --------------------------------------------------------------------------------------------- | ------------------ | ------------------------ | -------------------- | | 1 | Photo + keywords or short sentence → AI-extracted entry | Highest | High | Top | | 2 | Natural-language conversational query (single-turn or multi-turn) → trust-tagged result cards | High | High | Top | | 3 | Multi-language contribution and retrieval | High | High | Top | | 4 | Voice input with auto-transcription (both sides) | High | High | Strong (fast-follow) | | 5 | WhatsApp forward → AI-extracted entry | Medium-high | Medium | Strong | | 6 | Cross-friend gap-filler | Medium | Medium | Strong | | 7 | Geolocation-triggered prompt | High (if accepted) | Medium-low (permissions) | Later | | 8 | Annual memory book | Medium (retention) | High | Later | | 9 | Calendar-aware prompt | Medium | Low (permissions) | Later | | ||||
| DESIGN | Define Target State Workflow | UX Flows & Wireframes Suggested Tool: Excalidraw | Workflow (future) | Assuming your product or feature works as desired, what is the target state workflow? | Please leave this area blank. This space is for the Instructor to provide you with feedback. | |
| Build Wireframes | Wireframes | How will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features? | [Insert your response here. Include link to visuals as appropriate.] | |||
| Develop Prototype to showcase AI interactions | Prototype Screens Suggested Tool: lovable.dev | What aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases? | [Insert your response here. Include link to visuals / prototype as appropriate.] | |||
| Initial Prompt Design | Master Prompt [Initial Design] | Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency? | [Insert your response here] | |||
| Prepare for Testing & Iteration | Evaluation Criteria & Test Plan | Evaluation Criteria | What specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output? | [Insert your response here] | ||
| Example Cases | What specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs? | [Insert your response here] | ||||
| DEVELOP | AI Model Selection & Justification | AI Model Selection & Justification | Which AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product? | V1 uses two models, both reached over HTTP from server-only code so keys never touch the client. **`claude-sonnet-4-6` — contribution (extraction + compose).** One model handles every language task on the contribution side. It earns the slot on three properties V1 depends on: 1. **Multimodal vision.** The extraction call reads photo pixels to draft the note and to break ties among candidate places. The schema and prompt forbid it from setting place identity — vision is a signal, not the source of truth. 2. **Tool use for structured output.** Extraction is a forced tool call (`submit_recommendation`), so the draft arrives as validated JSON, not prose to parse. 3. **Instruction-following for the trust constraints.** The note must stay first-person, verbatim-faithful, and free of invented operational data. Sonnet holds these constraints across the extraction call and the two compose calls. The extraction call runs at temperature 0.2 with a cached system+tools prefix. `composeNote` runs at 0.3 (300-token cap), `composeTip` at 0.2 (80-token cap), in parallel. Every call has a deterministic fallback when no key is configured, so the app and the demo still run without Anthropic credentials. **`voyage-3.5` at 1024 dimensions — semantic search.** Anthropic has no embeddings endpoint, so embeddings use a separate Voyage key. Voyage was chosen for two reasons: `voyage-3.5` is multilingual, so the vectors stay valid for the V2 EN/FR/DE/LB corpus with no re-embed; and it supports asymmetric `input_type` (`document` for places, `query` for the Ask), which ranks short queries against longer place text well. The model and dimension are pinned in `lib/embed.ts` and must stay in lockstep with the `vector(1024)` column. **Cost and latency notes.** Contribution fires one extraction call at upload and two small compose calls at publish. Discovery fires one query embedding per search and one document embedding per publish/edit. The per-publish embed runs against Voyage's free-tier 3-requests-per-minute cap — fine at prototype scale, flagged for V2 rate-limiting. **Limitations and integration.** Sonnet's vision can over-read a photo, so the architecture treats vision as a signal only and never lets it set place identity. Embeddings are rate-limited on the free tier and the concatenated-note vector is unweighted (both logged for V2). Neither model sits on the trust-critical retrieval path — that is deterministic — so a model outage degrades polish and relevance, never trust. Both integrate as server-side calls behind Next.js route handlers (`/api/extract`, `/api/compose`) and a Supabase RPC (`match_places`); keys never reach the client. | Please leave this area blank. This space is for the Instructor to provide you with feedback. | |
| Define Inputs | Input Specification Table | Required Fields | What are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement. | **Contribution inputs.** - *Required:* one photo. - *Derived (client):* EXIF GPS latitude/longitude, when present. - *Resolved (server):* a Google Places result or a short candidate list, distance-ranked from the EXIF point. - *Optional (Review):* the parent's "what the child loved" text and "heads-up" text. The extraction call emits one `submit_recommendation` tool call with three sections: - **`place`** — `name`, `type`, `commune`, `country` (ISO-2), `cross_border`, `google_place_id`, `lat`, `lng`. Identity is trusted from Google Places; the model may only fill these, never invent a place. - **`recommendation`** — `note_seed` (≤ 2 sentences, plain parent voice), `tip` (nullable; photo-visible heads-up only), `age_bands` (conservative subset of `under 1` / `1-2` / `2-4` / `4+`, silent), `photo_note` (one high-confidence sentence naming what is visibly in the photo, else null), `expires_at` (ISO date, null = permanent default). - **`confidence`** — `place_match_score`, `geo_signal_score`, `candidate_count`, `overall` (= min of the two scores), `rationale`. Computed but not gating; not persisted. **Compose inputs** (`/api/compose`): `placeName`, `placeType`, `noteSeed`, `lovedAnswer`, `photoNote`, `headsUp`. Returns `{ note, tip }`. **Retrieval inputs** (`fetchDiscovery`): the query string, the authenticated user (resolved to their friend set and commune via RLS + the admin client), a sort key, and a `maxKm` (default 50). | ||
| Optional Fields | Are there any optional or user-customizable fields? How do they impact the AI’s output? | The only user-customizable inputs are the two Review free-text fields. The "what the child loved" answer is the spine of `composeNote` — it becomes the published note in the parent's own words; skip it and the flow falls back to the photo-only draft. The "heads-up" answer becomes the `tip`, tightened to one line; a decline stores no tip. Both are editable before publish, and the parent's edit is the final stored text. Nothing else is user-tunable — place identity, distance, and ranking are deterministic and out of the parent's hands by design. | ||||
| Define Good Output | Output Evaluation Checklist | Objective Criteria | What criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance) | **Objective criteria (machine-checkable).** The extraction tool call returns schema-valid JSON. Place identity comes from Google Places, never from the image. No operational field (hours, parking, café, stroller) is invented — the schema has no slot for one. The stored note and tip equal the parent's confirmed text, byte for byte. `age_bands` is a conservative subset, never "all ages" without support. `expires_at` is null unless a clear time-bound signal is present. On retrieval, every card resolves to a friend-visible place, distance is present and within 50 km, and the Maps URL resolves. | ||
| Subjective Criteria | Are there any criteria that require human judgment or qualitative assessment? | **Subjective criteria (human judgment).** Does the composed note read in the parent's own first-person voice rather than a brochure? Is the heads-up genuinely useful — the specific thing only a visitor would know? Does the recommendation meet the three-beat "great recommendation" standard? These need a human (and a human-calibrated LLM judge), because "sounds like a parent, not marketing" is not assertable in code. The binary graders and targets are in [[#Evaluation Criteria & Test Plan]]. | ||||
| Prompt Design Iteration | Master Prompt [Final Design] | Prompt Version 1 | What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints. | **Persona:** "Pebble's recommendation assistant" — a parent-friend voice, defined in the extraction system prompt below. - **Inputs:** the photo (multimodal), EXIF GPS, the resolved place or candidate list from Google Places, and any contributor text — passed in an XML `<contribution>` block. - **Initial instructions / constraints:** the eight rules below — identity from Places not pixels; no invented operational data; a grounded ≤ 2-sentence note seed; conservative age bands; null-default expiry; deterministic confidence; parent-not-marketing voice; high-confidence-only photo note. - **Variations tested:** answer chips on vs. off; `note_seed` vs. `photo_note` fed to compose; the ops-attestation block in vs. out; compose temperatures — all recorded in [[#Prompt Iterations & Evolution]]. - **Optimization techniques:** a forced tool call for a fixed output shape; prompt caching of the static system + tools prefix; low temperature (0.2) on the structured call; grounding constraints plus the vision-as-signal rule to suppress hallucination; deterministic fallbacks so a model failure never blocks publish. | ||
| Prompt Iterations | If revised, what changes did you make and why? How do you track and record prompt evolution? | The V1 prompts above are the product of several documented revisions during the build. Each change fixed an observed failure, not a hypothetical one: 1. **Generic-boilerplate note → conversational Review.** A static editable draft produced bland notes ("Zoo in Bettembourg. Worth a visit with the little ones.") that parents published unchanged. The fix was the four-turn micro-interview that pulls the parent's own words. 2. **Answer chips removed from the "loved" turn.** Photo-grounded answer chips duplicated the photo observation (`photo_note`) and added friction; an open text field replaced them. 3. **`note_seed` no longer fed to compose.** `note_seed` is a photo-based *guess* and was the source of fabrication ("peacocks roam free so your toddler can walk up to them"). Compose now receives `photo_note` — a factual observation — plus the parent's answer, and nothing else. 4. **Operational-attestation block removed from extraction.** The model no longer drafts hours/parking/café/stroller; the product defers to Google Maps. 5. **`composeTip` constrained to clean-and-shorten,** with a deterministic `meansNoHeadsUp` guard so a decline never becomes a fabricated heads-up. 6. **Temperatures tuned:** extraction 0.2 (structured), note compose 0.3 (voice), tip compose 0.2 (faithful). 7. **Confidence kept as a reading, not a gate** — the ≥ 0.90 publish gate was designed but not wired; disambiguation moved to Review Turn 0. Evolution is tracked in three places: the `pebble-conversational-review-spec.md` and `pebble-claude-code-brief.md` decision logs, the `CLAUDE.md` change log, and git history. | ||||
| Data Preparation & RAG Implementation | Data Preparation & RAG Implementation | What data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information? | Retrieval-augmented generation in V1 is retrieval-only: the corpus is the friends-only place set, and the "generation" is deterministic card rendering, not a model. **Embedding text recipe.** `placeEmbeddingText` builds one string per place: `"{name} — {type} in {commune}."` plus the concatenation of every recommendation note on that place. Name and type carry the topical signal ("indoor soft play", "café with kids' corner"); the notes add use-case language ("perfect for a rainy day") that lifts recall on natural queries. The live path and the backfill script share this one function, so a place embedded either way ranks identically. **Live re-embedding.** `publishRecommendation` calls `reembedPlaceFromNotes` after every insert — for new and deduped places — rebuilding the place vector from all its notes so a second friend's note reaches the vector immediately. `updateRecommendation` (note edit) does the same. It is best-effort: a Voyage hiccup never blocks publishing, and `npm run embed` backfills any gap. **Query path.** The Ask query is embedded with `input_type=query` and matched through `match_places(query_embedding, match_threshold, match_count=24)`, a `stable SECURITY INVOKER` SQL function returning visible places whose cosine similarity clears the threshold, best-first. `SECURITY INVOKER` is load-bearing — it keeps the friends-only RLS policy active inside the function. **Threshold and empty state.** `VECTOR_MATCH_THRESHOLD` defaults to 0.5 and is calibrated with `npm run embed:calibrate`. Nothing clearing the bar returns the empty state. Two known tradeoffs are logged for V2: one embed call fires per publish/edit (free-tier rate limit), and the concatenated-note vector is unweighted (note count and length both skew it). | |||
| Create Evaluation Set | Example Input/Output Data for Testing | Typical Examples | What are the most common inputs and expected outputs? Use real data if possible. | **Contribution, fast path.** Input: an EXIF-tagged photo of a single unambiguous farm. Output: place resolved from Places; Turn 0 confirms it; the parent types "she loved feeding the goats"; `composeNote` returns a first-person note weaving that with the visible detail; `tip` empty; `expires_at` null; one recommendation row with the parent's photo. | ||
| Edge Cases & Negative Cases | What examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain) | **Contribution, disambiguation.** EXIF lands between two candidate places within the ambiguity margin. Turn 0 presents the candidates; no auto-pick. Place identity never comes from the image. **Contribution, EXIF stripped** (screenshot or messaging re-encode). No GPS → `geo_signal_score` 0.0 → Turn 0 falls to "No / I'll add it" text entry; the system asks rather than guesses. **Retrieval, friend match.** "place with animals near a forest for my 2-year-old" from a user whose friend recommended such a place → one card, the friend's verbatim note, distance in km, Open-in-Maps. Several friends on the same place → one card stacking their notes and photos, most recent first. **Retrieval, no match.** A query that clears no place above the threshold, or matches only a non-friend's place, returns the canonical empty state — never a weak or non-friend result. | ||||
| Test Example Data & Review Results | Manual Review | Run your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why? | [Insert your response here] | |||
| Automated Evaluation | What pass/fail rate or scores did the AI achieve on core criteria? | **Status, stated honestly.** The deterministic assertion suite for the trust-critical units (friend filter, dedup, expiry, distance, Maps URL, place-resolution selection) is the standing CI gate and the layer the build relies on. The LLM-as-judge golden-set runner is *specified but not yet executable*: the 32 current test images are programmatically generated illustrations — valid as EXIF and ask-path fixtures, but not as vision fixtures — and are to be replaced with real geotagged photos per `PHOTO_SOURCING_GUIDE.md` before judge scores mean anything. No automated pass/fail percentages are claimed yet, by design rather than omission. | ||||
| Handle Edge Cases & Iterate | Edge Case Identification | What edge cases did you identify in testing or real usage? | **Edge cases surfaced in the build:** EXIF-stripped images (screenshots, messaging re-encodes) → place-name ask, never a guess; HEIC photos → converted client-side (`heic-to`) before upload; one place recommended by several friends → dedup to one card; cross-border places (Trier, Metz) → resolved and Maps-linked with no special-casing. | |||
| Updates & Adjustments | What prompt or system adjustments have you made based on failures, feedback, or edge case observations? | The adjustments are the seven revisions in [[#Prompt Iterations & Evolution]] — most directly the conversational Review, feeding `photo_note` (not `note_seed`) to compose, and adding the relevance threshold to retrieval. Each shipped in response to an observed failure above. | ||||
| Automate Evaluation Approach | Evaluation Method | What is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets? | Three methods, matched to the unit type and run at different cadences: - **Script (deterministic assertions)** for the trust-critical units. Runs on **every commit** in continuous integration; must be green to merge. Seconds to run, free. - **Model grader (LLM-as-judge), human-calibrated,** for the probabilistic contribution units over the frozen golden sets. Runs on **every prompt change and before each release**; a regression on any 100%-target grader blocks release. Each judge is validated against James's labels (≥ 90% agreement) before it gates anything. - **Human online review** runs **weekly during the 30-day prototype**: read every flagged card and empty-result query, open-code new failure modes, append regression cases. | |||
| Evaluation Frequency | How often will you re-run evaluations for new data, new prompts, or post-launch monitoring? | Every commit runs the assertion suite; every prompt change and pre-release runs the judge suite; the online review runs weekly during the prototype, as the cadences above specify. Scaling: golden sets are append-only (new cases added, never silently edited), and the lightweight harness is built to move to a hosted eval tool as volume grows. At 5–8 users no heavyweight platform is needed. | ||||
| DEPLOY | Finalize Launch & Rollout Plan | Operational Readiness Checklist | Technical Readiness | Is infra (APIs, databases, rate limits, monitoring, rollback) tested and documented? | **Infrastructure, rate limits, and rollback.** The external dependencies are four HTTP APIs: Anthropic (extraction + compose), Voyage (embeddings), Google Places (resolution), and Supabase (Postgres, auth, storage). The known rate constraint is Voyage's free-tier three requests per minute — fine at prototype scale, flagged for a paid tier on growth. Rollback is built into the platform: Vercel deploys are immutable and revert instantly to a prior build, and Supabase migrations are versioned and applied migration-first ahead of dependent code, so a schema change never half-lands. Monitoring at this scale is server logs plus the in-product flag and empty-state telemetry; no separate observability stack exists yet. | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Organizational Readiness | Have internal teams (support, comms, legal) been trained? Is documentation complete? | **Internal readiness and documentation.** Pebble is a solo-founder prototype, so "train support, comms, and legal teams" collapses to one person. Documentation is complete for that reality: the repo `README`, `AGENTS.md`, and `PLACE_RESOLUTION.md`; the `pebble-claude-code-brief.md` and `pebble-conversational-review-spec.md` build specs; and the `CLAUDE.md` project memory that carries every locked decision. | ||||
| Launch & Rollout Strategy | Launch Approach | What is your launch approach? Pilot, AB test, or all users—who gets access and when? | V1 launches as a closed, invite-only pilot to 5–8 friend users — not an A/B test, which is meaningless at this sample size, and not a public release. James seeds the first friend edges; each new user arrives through an existing friend's invite link, so the graph is connected from day one. Expansion past the pilot is gated on the trust-loop signal — do parents contribute, and do searches return trusted results — not on a date. | |||
| Scale Readiness | How will you ensure readiness for scale? How will you monitor initial volume and scale up? | Scale is not a V1 concern: Postgres with an HNSW pgvector index handles orders of magnitude more than the prototype corpus, and the binding growth constraint is the Voyage rate limit, whose mitigation (paid tier, batched embeds) is understood. Initial volume is watched through server logs and the metrics below. | ||||
| Go-to-Market Plan | Marketing / Training Assets | What assets (FAQ, demo, guides) will you prepare for external communication/marketing? | The primary go-to-market asset is the 4-minute capstone demo video, shot from the running product against a locked storyboard; it doubles as the explainer for new invitees. Beyond it, launch needs only an invite link and a one-line description — at 5–8 trusted users there is no FAQ, help center, or marketing site to build. | |||
| Stakeholder / Internal Comms | How will you communicate launch plans, progress, and outcomes internally? | Internal communication is not a multi-team concern for a solo founder: plans, progress, and outcomes live in `CLAUDE.md` and the capstone submission, and reach the prototype cohort over the channel they already use. | ||||
| Confirm Legal, Privacy & Risk Protocols | Data & Privacy | How do you handle and protect user data, including storage, privacy, and compliance? | [Insert your response here] | |||
| Policy & Compliance | Are content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain? | [Insert your response here] | ||||
| Define Success Metrics | Success Metrics | User/Business Metrics | What user metrics will indicate success? What business metrics will demonstrate value? | The user-success metrics are the three pillars below — accuracy, quality of reply, data availability — plus 3-month retention. The business metrics are the revenue diagnostics (free-to-premium conversion, ARPU, churn), tracked only post-product-market-fit. | ||
| AI Metrics | How will you measure AI performance and accuracy? | AI accuracy rolls into Pillar 1 for production (the "this looks wrong" rate plus the place-ID and expiry diagnostics) and into the Level-1/Level-2 evaluation suites pre-release (assertions + human-calibrated judges). The offline graders and the online flag rate are reconciled in the weekly review. | ||||
| Monitor, Iterate & Improve | User Support & Feedback Plan | Support Channels | Where can users get support? Is escalation and ownership clear? | The in-app "Something off?" flag is the V1 trust-repair channel. It captures one of four reasons (`wrong-age`, `closed`, `inaccurate`, `other`) and queues for offline review; there is no automated re-verification in V1. Gratitude ("Thank {name}") is the positive feedback loop and the prestige mechanic that earns titles like "Trusted voice in {commune}." At a cohort of 5–8 friends, James is the direct support line. | ||
| Feedback Workflow | How do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated? | **Feedback and bug triage.** Issues are triaged by trust impact, not volume. Trust-critical defects — a wrong-location publish or any friends-only filter leak — are P0 release-blockers, fixed before anything else and asserted against in continuous integration so they cannot regress. Flag reports are reviewed weekly and grouped by reason; a cluster in any reason becomes the next fix. Ownership is unambiguous at prototype scale — James triages, prioritizes, and ships — and a user's escalation path is a single message to him. | ||||
| Monitoring & Continuous Improvement | Monitoring Approach | What monitoring/logging is in place to spot operational/AI issues post-launch? | Three signals are watched weekly: flag volume by reason, the empty-state rate (the inverse of query coverage), and Open-in-Maps taps. One trigger is pre-decided — ship a structured age-band field only if `wrong-age` exceeds 10% of total flag volume after the 30-day prototype; below that, the silent inference is sufficient and the build is saved. The weekly loop reads every flagged card and empty-result query, open-codes new failure modes, and appends regression cases to the golden set (see [[#Evaluation Criteria & Test Plan]]). | |||
| Ongoing Improvement | How will you collect learnings, review performance, and update your system continuously post-launch? | The weekly error-analysis loop is the engine: read traces, open-code failures, append regression cases, fix, re-run the suites. Locked decisions land in `CLAUDE.md`; the golden set grows append-only so every past failure stays guarded. This is the same loop that drove the V1 prompt revisions, now run continuously. | ||||




