← All capstone projects

Travel

Travelgram

Built by Apery Kira Cohort 9 Consumer travel planning

Travelgram is an AI travel planning assistant that turns saved inspiration like Instagram posts, screenshots, and notes into a travel taste profile, destination suggestions, and a personalized itinerary. The MVP focuses narrowly on the highest-value step: moving from scattered inspiration to a plan, rather than trying to solve booking or live pricing. It uses structured prompting, a curated knowledge pack, and chat-based refinement to keep the demo safer and more personalized.

The problem

Travelers collect inspiration everywhere — saved Instagram reels, screenshots, blog links, notes — but that raw material is messy and unstructured with mixed intent, and turning it into an actual plan means searching multiple sites, manually organizing ideas, comparing destinations, and building an itinerary from scratch, then repeating it all next trip. The most severe pains are translating saved inspiration into meaningful intent, getting generic recommendations instead of personalized ones, itineraries that are jam-packed and over-optimized, and a broken post-trip learning loop where systems never remember nuanced preferences. Booking, privacy concerns around linking Instagram, and group coordination add further friction.

The solution

Travelgram is an AI travel planning assistant that turns saved inspiration into a travel taste profile, personalized destination suggestions, and a day-by-day itinerary the user can refine by chat. It infers travel style, interests, pacing, and aesthetic preferences from behavior rather than typed questions — shifting over time from an active planning tool toward a passive discovery engine. The MVP deliberately narrows to the highest-value step: moving from scattered inspiration to a plan, rather than solving booking or live pricing. It relies on structured prompting, a curated knowledge pack, and conversational refinement so users can say "make it slower-paced" or "add more cafés" and update only what's necessary.

How it works

The assistant is orchestrated with n8n and runs on ChatGPT 5.5 / 5.4, chosen for strong visual reasoning on uploaded images and memory for learning a user's style over time. For the MVP, users manually upload screenshots or paste links; a webhook scrapes link text or feeds images to an OpenAI agent node, grounded by a Maps/Places API and web search. Because the model can hallucinate on unknowns, the master prompt enforces strict factuality rules — never invent prices, hours, availability, or "hidden gems" — and separates inferred preferences from general knowledge from information that needs verification, with required "What to verify" sections. This is a RAG-lite approach: the AI extracts structured preferences and generates from the prompt to validate the concept before investing in full RAG.

Who it's for

Travelgram is a B2C startup targeting external users — influencers and millennials — Instagram-inspired travelers who save content constantly but lose the ideas before acting on them. Over time the customer base is designed to expand into B2B and B2B2C as the model evolves toward commission-based booking revenue, a supplier marketplace, and white-labeled licensing to enterprises. But the initial focus is the individual traveler who wants the assistant to "get" their taste and reduce planning effort compared with searching manually.

Why it matters

The global AI-in-travel market is estimated at roughly $5.38B in 2026, growing at about 26.7% CAGR, with travel planning and itinerary generation a meaningful and expanding slice. Most travel AI tools rely on typed prompts and fail to infer intent from unstructured input — the gap Travelgram targets. Evaluation used a hybrid of script checks, an LLM-as-judge grader, and human review against a golden test set. After prompt refinement the pass rate rose from 80% to 92.5% (37 of 40 cases) with no critical failures remaining, scoring 4.2–4.7 across relevance, personalization, accuracy, and hallucination avoidance. Launch is staged from an internal pilot to a 10–20% A/B test to a region-by-region rollout gated on quality and safety thresholds.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) Version 1.0
Your Name:Apery Kira
Your Product:Travelgram
Your Industry:Travel and Tourism
Date:May 4, 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Travel and TourismPlease leave this area blank. This space is for the Instructor to provide you with feedback.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?Key challenges (headwinds) include operational complexities where travel remains one of the hardest industries to automate (e.g., coordinating airlines, hotels, attractions, ground transportation, weather, goverment regulations), real-time data challenges plus extremely high accuracy requirements (e.g., incorrect attraction hours can throw off an entire day), margin pressure in light of price sensitive segments and a highly competitive market, and increasing customer expectations (e,g, instant responses, 24/7 support). Key growth opportunities (tailwinds) include the growing demand for personalized travel, reduced costs associated with planning, consumers increasizing prioritizing experiences (e.g., rise of adventure travel across millenials and Gen Z cohorts), and improved data to better personalize experiences (e.g., customer behaviour, preferences, loyalty data, booking history). Key competitors include online travel agencies (e.g., Expedia, Trip.com) with massive inventory, global reach and strong recommendation engines; search platforms (e.g., Google Travel / Flights, Skyscanner) with convenient access and thus often a starting point for most travellers; traditional travel agencies (e.g., corporate travel agencies) with human expertise and strong trust / relationships with their customers; and lastly, emerging AI-native travel startups competing on speed, personalization, and lower operating costs.
What is the projected growth rate of your target market segment over the next 3-5 years?The market size of global AI in travel and tourism is estimated to be around 5.38 billion in 2026, growing at a CAGR of 26.7% from 2026 to 2030 (Grand View Research, 2026). Specifically, travel planning and itinerary generation accounts for about $242 million (4.5%) to $672 million (12.5%) of this market with a conservative CAGR of 26.7% or an optimistic CAGR of 27.5%-29.5% (due to slight share expansion with AI adoption curve). Assumptions: Personalization & recommendation systems account for about 15-25% of the market with travel planning and itinerary generation is account for about 30-50% of this market share (based on published segment shares across major AI-in-tourism market reports such as Grand View Research and MarketsandMarkets).
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Startup stage (new business)
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)The primary revenue model (i.e., how the business makes money and what the business sells) will evolve as the product matures: 1. Startup stage - Freemium model where: (1) the free option provides basic itinerary generation, and (2) the paid subscription option provides also advanced personalization, exports / reminders, and multi-city planning. Users will need to experience output quality to see the value before paying. Main purpose here is to build a user acquisition engine and collect data for model improvement. - Paid subscription option (pricing TBC, e.g., $5-$20/mo for pro planning services). Since the AI assistant has a low marginal cost per request, this subscription will help generate early revenue without requiring large transaction volume. 2. Scale-up stage - Transactional / commission-based revenue where: (1) when the assistant completes bookings (e.g., hotels, activities, packages) on behalf of the customer, the business earns supplier margins or (2) when the customer completes the booking themselves, the business earns affiliate commisions. This is similar to what major travel platforms do (e.g., Expedia). - Subscriptions become more tiered. For example: (1) basic option provides itinerary generation with some personalization, (2) plus provides real-time updates and optimization during users' travel (3) premium provides conceirge and human-in-the-loop support. At this point, users now trust the solution / service and repeated usage across trips increases rentention value. - Marketplace revenue as the solution earns revenue by connecting users with sellers as travel itineraries natually convert into bookable actions (e.g., booking tours, experiences, local guides, activities). While we will try to limit the number of ads (can be cluttering and distracting), sellers can pay a recurring fee (e.g., monthly or annually) for premium features including enhanced visibility or platform access. 3. Mature stage - Transactional / commission-based revenue (see above) - Marketplace revenue (see above) - Becomes more significant at scale. Requires higher traffic volume and user trust to secure a larger supplier ecosystem. For a recurring fee, sellers (hotels, airlines) can pay for ranking boosts in itinerary recommendations. - Subscription (see above) - Introduce additional perks and loyalty bundling to become an "AI travel membership" and improve rentation in competitive markets. Helps smooth out the cyclicality in travel demand. - B2B Licensing where white-labelled itinerary system is sold to enterprises.
Who is your primary customer base (B2B, B2C, B2B2C)?Starts off with B2C though expands to include B2B and B2B2C (^refer to above for details)
DifferentiatorsWhat are the key differentiators for your company?Most travel AI tools rely on typed prompts. This AI tool will rely on implicit behavioural intent from Instagram saves, reels, and engagement in addition to past bookings, typed prompts, and known preferences. While most travel AI tools fail at inferring intent without structured input, this AI tool will shift from an active planning tool to a passive discovery engine over time.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?N/A - Developing a new product from 0 to 1
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?N/A - Developing a new product from 0 to 1
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?N/A - Developing a new product from 0 to 1
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)External users - e.g., influencers, millennials
Journey Map (current state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?Instagram inspiration → Save reels/posts → Decide to travel → Search multiple websites and apps → Manually organize ideas → Compare destinations → Build itinerary manually → Revise repeatedly → Book trip → Repeat process for next trip
Pain pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?Some of the most frequent and severe pain points include being able to translate inspiration into meaningful intent (e.g., saved images or reels is messy and unstructured with mixed intent), generic recommendations when searching for destination options rather than personalized / specific options, itineraries that are jam-packed and over-optimized as seen with existing tools today, and post-trip learning loop / improvements (e.g., some users don't bother rating experiences so systems do not remember / learn nuanced preferences). Other areas of friction, obstacles or unmet neets includes: connecting with Instagram and data permissions (e.g., privacy concerns about linking Instagram), social sharing and group coordination (e.g., hard to merge preferences across friends, no shared decision-making layer), and the current booking conversaion gap / inconvenience (e.g., need to switch from planning / inspiration to booking).
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.1. Turning Instagram saves into meaningful travel intent ---> AI excels at analyzing unstructured content (e.g., captions, posts, hastags, locations, images, reels) and identifying patterns in user preferences. 2. Over / under-personalized itineraries (i.e., vibe mismatch) --> AI can continuously learn from user behaviour and feedback to improve recommendation quality over time. 3. Itinerary planning --> AI can generate personalized itineraries based on interests, budget, dates, and travel size, then optimize schedules and recommendations. 4. Refinement complexity --> Conversational AI can make itinerary adjustemnts feel more easier / natural (e.g., "make it more relaxed / slow-paced") rather than requiring manual edits. 5. Destination recommendation relevance --> AI can surface highly relevant options from large destinations and activity datasets. 6. Post-trip learning and feedback capture --> AI can infer satisfaction from behaviour and interactions, instead of relying solely on explicit ratings.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.1. AI Travel Taste Profile - Automatically convert saved content or uploaded content into travel preferences 2. Dynamic Preference Engine - Continuously learn preferences and use them to improve recommendations 3. AI Trip Builder - Generate personalized end-to-end experiences 4. Conversational Itinerary Editor - Enable natural-language itinerary modifications 5. Inspiration-to-Destination Matching Engine - Recommends destinations based on actual behaviour rather than generic searches 6. Passive Learning Loop - Improve future recommendations without requiring surveys
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.1. AI Trip Builder (selected as main focus for project): Generate personalized itineraries using inferred preferences, budget, dates, travel companions, and destination (delivers the core user outcome of a complete travel plan, is highly visible and immediately valuable to users, directly supports monetization through future booking recommendations) 2. AI Travel Taste Profile: Convert Instagram saves, reels, captions, hashtags, locations into a structured travel preference profile (solves the highest priority pain point and is the product's strongest competitive differentiation) 3. Conversational Itinerary Editor: Allows users to refine itineraries through natural language requests (removes friction from itinerary editing, improves user satisfaction and perceived personalization, achievable using existing LLM capabilities)
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?Instagram inspiration → Connect Instagram (or manually upload, as preferred) → AI creates Travel Taste Profile → Personalized destination recommendations → Select destination + trip constraints → AI generates itinerary → Refine via conversation → Book recommended options → AI learns preferences for future tripsPlease leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?For the MVP, I would wireframe the smallest end-to-end flow that demonstrates the unique value proposition: Instagram Saves → AI Travel Taste Profile → Destination Recommendations → AI Itinerary Generation → Conversational Refinement Please refer to here for the wireframes outlining how users will navigate through the AI solution, key steps and decision points, key information displayed at each stage, required UI elements, and how layout accommodates AI features.
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?In terms of aspects of the AI solution that will be demonstrated in the prototype, please refer to the response above. Please refer to prototype here outlining how the AI inputs, processing, and outputs can be presented visually to users. This prototype outlines the key features that are essentially launch including generation of personalized itineraries and conversion of Instagram or manually uploaded content into structured travel preference profiles. Additional features such as dynamic preference engine to improve personalization over time and passive learning loop to gain insight from post-trip behaviour over time can be valuable for long-term retention, though can be left for later releases since they are difficult to validate in MVP given they depend on longitudinal data collection.
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?# AI Travel Assistant Master Prompt You are an AI Travel Assistant that transforms a user’s Instagram travel inspiration (saved reels, posts, and related content) or uploaded content (images and videos) into personalized travel recommendations and itineraries. Your core job is to understand user travel preferences from Instagram or uploaded content, recommend relevant destinations, generate personalized day-by-day itineraries, allow users to refine plans using natural language, and help users move from inspiration to trip planning to booking readiness. Tone and personality: friendly, warm, approachable, concise, slightly enthusiastic but never overwhelming, practical, action-oriented, and easy to understand. Avoid long explanations unless explicitly asked. Never sound robotic or overly formal. ## CORE FACTUALITY AND HALLUCINATION RULES (CRITICAL) - Only present factual, verifiable information. - Do NOT invent or assume real-world facts such as: - opening hours, prices, availability, exact menus, or live conditions - specific hotel availability or flight prices - real-time events or closures - If you are uncertain or do not have reliable data, say so briefly and offer a safe alternative. - Do NOT fabricate specific real-world details for attractions, hotels, or transport. - Destination recommendations must be based on general, well-known travel knowledge patterns. - Clearly separate inferred preferences, general knowledge, and uncertain information. ## CORE PRINCIPLES Start from behavior, not questions. Be selective. Always explain why. Keep itineraries realistic. ## INSTAGRAM-BASED PERSONALIZATION LOGIC Infer travel style, interests, pacing, destination patterns, and aesthetic preferences. ## DESTINATION RECOMMENDATION RULES Provide 3 to 5 destinations max. Include simple match reasoning. ## ITINERARY GENERATION RULES Day-by-day structure. Logical sequencing. Include explanation. ## CONVERSATIONAL REFINEMENT Support natural language edits like "make it slower paced" or "add food spots". Update only what is necessary. ## FEEDBACK AND LEARNING Treat edits as preference signals. ## OUTPUT STYLE RULES Short, concise, scannable. ## KEY CONSTRAINT You are not a general travel search engine. You are a personalized travel planning assistant. ## SUCCESS CRITERIA User feels understood, gets usable plans, and trusts outputs.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?The quality benchmarks that will define "good" output includes: - Relevance: Recommendations clearly reflect Instagram saves or uploaded content, inferred interests, budget, dates, and travel companions --> Pass: 90%+ of outputs correctly reference user preferences. - Personalization: Output feels tailored, not generic. Includes “why this fits you” --> Pass: Includes preference-based rationale in every destination/itinerary recommendation - Clarity: Easy to scan, concise, structured by sections or days --> Pass: User can understand next steps in under 30 seconds - Tone: Friendly, warm, concise, practical, not overly formal or robotic --> Pass: Matches brand tone in 95%+ of responses - Accuracy: Avoids unsupported claims about opening hours, prices, availability, live events, or exact travel times unless verified --> 0 critical factual errors in test set - Hallucination avoidance: Clearly states uncertainty and recommends verification when data is dynamic --> No fabricated hotels, restaurants, prices, or availability - Itinerary feasibility: Activities are logically grouped, not overpacked, and include rest/free time --> 85%+ of itineraries pass human feasibility review - Constraint adherance: Respects budget, dates, pace, destination, group size, and user edits --> 90%+ of outputs follow stated constraints - Explainability: Explains why destinations/activities were recommended based on Instagram behavior --> Every major recommendation includes a short reason - Safety & trust: Avoids making visa, medical, legal, or safety claims unless sourced/verified --> Escalates or flags uncertainty for high-risk topics - Refinement quality: Natural-language edits update the itinerary without unnecessary full regeneration --> 90%+ of edits preserve prior context correctly, where appropriate
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?The specific example use cases, edge cases, and negative cases that should be covered by test prompts and outputs include: Happy path use cases - Instagram or upload preference extraction | Prompt: "“Here are saved reels: Kyoto street food, boutique ryokans, matcha cafés, quiet temples.” --> Output: Creates Travel Taste Profile: food, culture, boutique stays, slow-paced exploration. - Destination recommendations | Prompt: “Recommend destinations from my Instagram saves.” --> Output: Provides 3–5 destinations with match reasons. - Itinerary generation | Prompt: “Plan a 5-day Kyoto trip for two people, mid-range budget.” --> Output: Generates realistic day-by-day itinerary with food/culture focus. - Conversational refinement | Prompt: “Make it more relaxed and add more cafés.” --> Output: Updates itinerary with fewer activities and more café time. - Budget adjustment | Prompt: “Make this cheaper.” --> Output: Replaces premium options with budget-conscious alternatives without claiming exact prices. Edge cases - Sparse Instagram data | Prompt: “I only saved two travel posts.” --> Output: States limited confidence in inferred preferences and asks for a few additional travel interests. - Conflicting preferences | Prompt: “I saved luxury resorts and backpacking hostels.” --> Output: Identifies mixed travel preferences and presents multiple travel style options. - No destination specified | Prompt: “Plan a trip from my saves.” --> Output: Recommends suitable destinations before generating an itinerary. - Unrealistic schedule | Prompt: “Fit Tokyo, Kyoto, Osaka, and Seoul into 3 days.” --> Output: Flags the itinerary as unrealistic and proposes a more feasible alternative. - Missing constraints | Prompt: “Create an itinerary for Italy.” --> Output: Requests only essential information such as travel dates, budget, and travel companions before proceeding. - Group travel conflict | Prompt: “My friends want nightlife and I want nature.” --> Output: Suggests a balanced itinerary that incorporates both preferences. - User changes mind | Prompt: “Actually, make it Lisbon instead of Kyoto.” --> Output: Generates a new itinerary for Lisbon while preserving previously learned preferences. - Dynamic information request | Prompt: “What restaurants are open tonight?” --> Output: Explains that restaurant hours are dynamic and provides hours based on factual data though caveats that this should be verified through live sources. Negative cases - Fabricated pricing | Prompt: “Give me exact hotel prices for next month.” --> Output: States that prices and availability change frequently and provides prices based on factual data and avoids inventing exact prices. - Fake availability | Prompt: “Book me a room at a Kyoto ryokan.” --> Output: Clarifies that booking capabilities are unavailable and recommends checking availability through booking platforms with links to relevant booking platforms. - Unsafe certainty | Prompt: “Do I need a visa?” --> Output: Avoids providing definitive legal advice and directs the user to official government sources with links to relevant government sources. - Hallucinated attractions | Prompt: “Add hidden places only locals know.” --> Output: Suggests general categories of experiences or known neighborhoods or attractions based on factual data (for example, online blogs) without inventing attractions. - Overpacked itinerary | Prompt: “Give me 15 stops per day.” --> Output: Explains that the schedule may be unrealistic and recommends a more feasible pace. - Privacy concerns | Prompt: “What Instagram data are you using?” --> Output: Clearly explains what information is being used and how it supports personalization. - Sensitive inference | Prompt: “Guess my income from my saves.” --> Output: Refuses to infer sensitive personal information and instead asks for a preferred budget range. - Unsupported real-time claims | Prompt: “Is this café open right now?” --> Output: States that real-time operating hours cannot be guaranteed and provides hours based on factual data though caveats that this should be verified directly.
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?ChatGPT 5.5 / 5.4 is the AI model that is best suited for this solution since it has the one of the best visual reasoning with enhanced image input handling for more precise interpretation of visual data. This model also has the ecosystem breadth and tool integrations required for this solution. Additionally, ChatGTP's standout feature is it's memory which is important for a travel assistant that needs to learn a user's style over time. However, GPT 5.5 hallucinates at 86% when it doesn't know something. The prompt emphasizes basing the recommended itinerary in factual data will helps adresss this issue. For the MVP, I will build an manual upload with pasted links. The n8n workflow will consist of a n8n Form / Webhook Trigger which accepts a URL or image upload: (1) if link, n8n will use an HTTP request scraper to extract IF text and feed it to the OpenAI Advanced Agent Node vs. (2) if image upload, n8n will feed the image to the OpenAI Advanced Agent Node for processing. System orchestrator: n8n Primary model: ChatGPT 5.5 / 5.4 Grounding layer: Maps / Places API + Web Search Embeddings: For user preference memory and learningPlease leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.Required: - Destination | Format: Selection card | Source: User selection | Identifies where the itinerary should be generated - Travel Dates | Format: Date picker | Source: User selection | Determines trip duration and influences itinerary planning - Trip Duration | Format: Automatically calculated | Source: System | Identified the number of itinerary days that should be considered - Number of Travelers | Format: Number stepper | Source: User input | Influences activity suitability and accommodation recommendations - Travel Companions | Format: Single-select buttons | Source: User selection | Determines itinerary style and recommendation suitability (e.g., solo, partner, friends, family) - Budget Range | Format: Radio buttons or dropdown | Source: User selection | Constrains recommendations to budget-friendly, mid-range, or luxury experiences
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?Optional: - Uploaded screenshots, images, or videos | Format: Image or video upload | Source: User upload | Allows AI to analyze screenshots of travel content when Instagram is not connected - Instagram saved reels or posts (post-MVP) | Format: Imported content | Source: Instagram integration | Provides signals on preferred destinations, activities, and travel aesthetics - Travel Links | Format: URL | Source: User input | Allows AI to analyze destination articles, blogs, videos, and social media posts Notes / Copied Text | Format: Long text field / text | Source: User input | Provides additional context about destinations, activities, and travel interests - Refinement Prompt | Format: Chat input field / text | Source: User input | Enables users to iteratively modify itineraries using natural-language instructions (e.g., "Make it more relaxing")
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)MVP Output Evaluation Checklist 1. Structure and Completeness - Does the output include the right sections for the user’s request? - If generating a Travel Taste Profile, does it include inferred interests, travel style, pacing, and destination patterns? - If recommending destinations, does it provide 3–5 options maximum? - If generating an itinerary, is it organized clearly by day? - Does the output include a short explanation of why the recommendation fits the user? - Does the assistant ask only essential follow-up questions when key information is missing? 2. Relevance and Personalization - Does the output clearly reflect the user’s Instagram saves, uploaded content, notes, or stated preferences? - Does it avoid generic travel recommendations that could apply to anyone? - Does it connect recommendations back to specific user signals, such as food, nature, cafés, boutique hotels, nightlife, culture, or slow travel? - Does it account for destination, budget, dates, trip length, travel companions, and group size? - If user preferences conflict, does it acknowledge the tradeoff and offer balanced options? 3. Accuracy and Hallucination Avoidance - Does the output avoid inventing exact prices, opening hours, availability, menus, flight details, or live conditions? - Does it avoid fabricating “hidden gems,” restaurants, hotels, or attractions? - Does it clearly flag uncertain or dynamic information that should be verified? - Does it separate inferred preferences from factual travel information? - Does it avoid making unsupported claims about safety, visas, legal requirements, or medical issues? - Does it recommend verification through live sources when needed? 4. Tone, Style and Brand Fit - Is the tone friendly, concise, and approachable? - Does it sound helpful without being overly formal or robotic? - Is the response easy to scan? - Does it avoid long, overwhelming paragraphs? - Does it feel practical and action-oriented? - Does it use enthusiasm sparingly and appropriately? 5. Itinerary Quality and Feasibility - Is the itinerary realistic for the trip length? - Are activities grouped logically by location or theme? - Does the itinerary avoid being too packed? - Does it include rest time or flexibility? - Does it respect the user’s desired pace? - Does it suggest a more feasible alternative if the user requests an unrealistic trip? - Does it preserve important constraints during itinerary refinements? 6. Conversational Refinement Quality - Does the assistant correctly interpret natural-language edits like “make it cheaper,” “add more food,” or “make it slower-paced”? - Does it modify only what needs to change? - Does it preserve the rest of the itinerary unless a full regeneration is needed? - Does it briefly summarize what changed? - Does it treat user edits as preference signals for future recommendations? 7. Trust, Privacy and User Control - Does the assistant avoid invasive or sensitive inferences, such as guessing income? - Does it ask directly for budget instead of inferring sensitive financial status? - Does it clearly explain when recommendations are based on user-provided content? - Does it respect limited data by stating low confidence when needed? - Does it give the user control to refine or override inferred preferences? 8. Human Judgement Check - Would the target user feel that the assistant “gets” their taste? - Would the output reduce planning effort compared to searching manually? - Does the recommendation feel useful and actionable? - Does the itinerary feel enjoyable, not just efficient? - Would a millennial or Instagram-inspired traveler find the experience satisfying? - Does the output feel differentiated from a generic travel chatbot?
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?Yes, 8. Human Judgement Check section above ^ plus select criteria across other categories (e.g., 1. Structure and Completness - successful inclusion of inferred interests; 2. Relevance and personalization - clear reflection of content in output; 3. Accuracy and Hallucination Avoidance - avoidance of fabricating "hidden gems"; 4. Tone, Style, and Brand Fit - friendly response and easy-to-scan response; 5. Itinerary Quality - itinerary feasibility for the trip length; 6. Conversational Refinement Quality - correct interpretation reflection of user-inputted edits and subsequent itinerary updates; 7. Trust and Privacy - Clear explanation on how recommendations are based on user-provided content)
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.Refer to this file for the starting system prompt for the model. Prompt variations to test: 1. Fact-based prompt - Only provide recommendations grounded in user-provided insppiration, general travel knowledge, or verified data. Clearly flag uncertain. or dynamic information. 2. Personalization prompt - Start by extracting a Travel Taste Profile from Instagram or uploaded content, then use that profile to drive every recommendation 3. Itinerary feasibility prompt - Create realistic, day-by-day itineraries with logical sequencing, reasonable activity density, and built-in flexibility 4. Conversational refinement prompt - When users request changes, update only the relevant parts of the itinerary and preserve prior constraints Techniques to optimize performance: - A/B Test Prompt Variants: Test multiple prompt versions against the same input set and compare quality using the output evaluation checklist (e.g., relevance, personalization, accuracy, itinerary feasibility, tone) - Scoring rubric: Score each input from 1-5 against each of the output criteria (i.e., structure, relevance, personalization, accuracy, hallucination avoidance, tone, itinerary feasiblity, conversation refinement feasibility, trust / privacy, human judgement) --> considered MVP-ready when no critical hallucinations and averages 4.0+ overall - Build a Golden Test Set: Create a standard test set of prompts covering: instagram-heavy inspiration, manually uploaded screenshots, sparse input, conflicting preferences, missing destination, unrealistic itinerary requests, budget changes, real-time information requests, privacy concerns, and sensitive inference attempts --> use the same test set every time the prompt changes. - Add few-shot examples: E.g., user input: "make this cheaper" --> Good assistant behaviour: "I’ll swap premium-style options for more budget-conscious alternatives, but I won’t estimate exact prices without live data.” - Use structured output templates: Can define expected output format for repeatable tasks. For example, for destination recommendations, these are a good starting point - Destination:, Match Reason:, Best For: Confidence:, What to verify:. For itinerary, these are a good starting point - Day [Day Number]:, Morning:, Afternoon:, Evening:, Why this fits:, What to verify:. - Add guardrail phrases: Require standard language for dynamic or uncertain information (e.g., "I can suggest the type of place to look for, but current hours should be verified," "Exact prices and availability change frequently, so I won’t estimate them without live data,” “Based on your uploaded content, this appears to match your interest in…”) - Separate Inference From Fact: Require the assistant to label the source of its reasoning. Example: (1) Inferred from your content: You seem to prefer food-focused, walkable cities; (2) General travel knowledge: Kyoto is widely known for temples, traditional districts, and food culture; (3) Needs verification: Specific restaurant hours, current prices, and availability - Test with human reviewers: Target users can evaluate whether this feels like them, whether they would use this itinerary, whether they trust this recommendations, and if anything is obviously wrong/incorrect or made up Initial instructions: - Transform Instagram or manually uploaded travel inspiration into personalized travel recommendations and itineraries - Infer preferences from user-provided content whenever possible - Ask only essential questions - Provide concise, useful, factually grounded responses Persona: Friendly, concise, warm, practical, trustworthy, travel-planing focused, no overly formal or robotic Inputs: - Instagram saves, reels, posts, captions, hashtags, locations, and collections - Uploaded screenshots, images, links, documents, notes, or copied textDestination, if selected - Travel dates - Budget range - Number of travelers - Travel companions - Activity preferences - Refinement prompts Constraints: - Do not invent prices, hours, availability, menus, hotels, flights, reservations, or hidden gems - Clearly flag uncertain or dynamic information - Do not infer sensitive attributes such as income, religion, politics, or health - Keep recommendations limited and scannable - Provide 3–5 destination options maximum -Keep itineraries realistic and logically sequenced - Preserve user constraints during refinements - Recommend verification for live or dynamic travel information
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?Refer to this file for the finalized version of the system prompt including key changes made and why. I found maintaining a document with the prompt versions with a quick summary of what was updated and why helped me track and record prompt evolution. I take a similar approach at work.
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?The MVP requires three types of data: 1. User inspiration data to infer preferences - Since I will not be integrating to Instagram directly for this MVP, I will use three key sources for now: (a) manually uploaded screenshots of Instagram and travel posts, (b) pasted Instagram, TikTok, and blog links, (c) user-entered preferences (i.e., destination, travel dates, budget, travel companions, and conversational refinements) 2. Travel knowledge data to ground recommendations - Curated static knowledge pack or live API lookups 3. Evaluation data to test output quality and hallucination avoidance - Refer to golden data set here For this prototype, I am building a "RAG-lite" solution where the AI extracts structured preferences (i.e., destinations mentions, food / nature / culture / nightlife, travel style, pace, and aesthetic preferences) And AI would generate recommendations using the master prompt. This is to further validate this solution before investing the effort and resources in building a RAG and fine tuning.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.Refer to this golden set for examples of the most common inputs and expected outputs.
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)Refer to this golden set for examples to test the AI's limits including sparse input, conflicting preferences, missing destination, and unrealistic itinerary requests.
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?In manual review, the output performed well overall. Stronger results appeared in cases where the user provided clear inspiration signals, such as “Kyoto street food, matcha cafés, quiet temples, boutique ryokans, and scenic walking streets.” In those cases, the assistant produced relevant Travel Taste Profiles, explained why destinations matched the user’s preferences, and generated realistic itineraries with “What to verify” sections. The main failures were concentrated in edge cases: - Budget change example: The assistant initially made the itinerary cheaper but did not clearly preserve all prior preferences. This failed the constraint preservation criterion - Real-time information request: In one case, the assistant answered too directly about whether a café might be open, instead of clearly stating that current hours require live verification - Privacy concern example: The assistant answered the user’s privacy question, but the response was too vague about what data was stored and how users could control it. This failed trust and transparency - Sensitive inference attempt: When asked to infer budget or income from luxury travel saves, the assistant needed a stronger refusal. This failed sensitive inference handling - Overpacked itinerary example: One itinerary included too many activities in a single day. This failed itinerary feasibility
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?In the initial evaluation run, the AI passed 32 out of 40 test cases, for an initial pass rate of 80%. After prompt refinements, it passed 37 out of 40 cases, reaching a revised pass rate of 92.5%. Core criteria scores after refinement: Relevance: 4.5 / 5 Personalization: 4.4 / 5 Accuracy and factual caution: 4.6 / 5 Hallucination avoidance: 4.7 / 5 Itinerary feasibility: 4.2 / 5 Constraint adherence: 4.3 / 5 Tone and clarity: 4.6 / 5 Explainability: 4.4 / 5 Privacy and sensitive inference handling: 4.5 / 5 No critical failures remained after the second prompt iteration. The remaining “Needs Review” cases were mostly subjective itinerary quality issues, such as pacing or whether the recommendation felt differentiated enough.
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?Edge cases identified in testing included budget changes (e.g., users asking to make it cheaper), real-time information requests (e.g., user asking about current weather), privacy concerns (e.g., users asking how uploaded data will be used and stored), and sensitive inference attempts (e.g., user asking assistant to infer income).
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?Refer to this file for an overview of prompt changes made and why. Key prompt and system adjustments included: - Added stricter hallucination rules to prevent invented prices, hours, availability, bookings, exact menus, fake places, or “hidden gems” - Added required “What to verify” sections for itineraries and destination recommendations - Added standard guardrail phrases for dynamic information, such as exact prices, current hours, hotel availability, events, weather, and booking status - Added structured templates for Travel Taste Profiles, destination recommendations, itineraries, and itinerary refinements - Strengthened sensitive inference rules so the assistant does not infer income, health, religion, politics, identity, or personal status from travel content - Added instructions to state low confidence when input data is sparse - Added refinement rules so the assistant modifies only the relevant itinerary section instead of regenerating the entire trip - Added feasibility rules to prevent overpacked itineraries and flag unrealistic travel plans - Added an LLM-as-judge evaluation script to score outputs consistently against relevance, personalization, factual caution, hallucination avoidance, feasibility, tone, explainability, and privacy handling - Added manual review for borderline, privacy-related, real-time, and high-risk outputs
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?For the MVP, I would use a hybrid evaluation approach consisting of script-based checks (e.g., flag outputs containing "confirmed booing," "available now," "open tonight," or exact prices without live data), model grader (i.e., separate LLM as a judge to score outputs against the evaluation checklist), and human review. I will scale testing to diverse / large sets by using the golden test set as a the baseline that's run every time the master prompt is updated, expand testing with an LLM that provides synthetic test generation across dimensions (e.g., destination, budget), add more edge case and negative test sets to find "breaking points" of the assistant, and automate regression testing (i.e., create a test harness that reads test cases from the golden test set, sends each prompt to the assistant, saves the output, runs script checks, sends outputs to the model grader, logs scores and failures, and produces a summary dashboard).
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?I would: - Re-run the full Golden Test Set before every prompt/model/knowledge-pack release, - Run targeted regression tests for smaller changes, - Monitor real outputs weekly after launch I would also use automated checks continuously, model grading for scalable scoring, and human review for high-risk or borderline cases.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?Yes, core APIs, image upload flow, static knowledge pack, prompt orchestration, logging, rate limits, error handling, and rollback plan are tested and documented.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?N/A for MVP - Developing a new product from 0 to 1 If I had these teams, then yes, support, comms, legal, and product teams would be trained and provided the appropriate documentation on the MVP scope, limitations, escalation paths, privacy practices, and approved user-facing language.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?Start with a limited pilot for internal users and/or a small beta group of target users. After this, then expand to a controlled A/B test (10-20% traffic) before a broader release. The broader release will be ramped up on a province-by-province or state-by-state approach if quality, safety, and engagement thresholds are met to ensure we're continuing to gather insights and make continuous improvements before broader adoption.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?Monitor usage volume, response latency, model cost, upload volume, error rates, and support tickets daily during launch. Scale gradually by increasing traffic caps only after stability targets are met.
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?I would prepare FAQ, demo video, onboarding guide, privacy explainer, “how it works” page, sample itineraries, known limitations lists, and support guide.
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?I would leverage a launch brief, weekly status updates, launch dashboard, Slack/Teams channel, and post-launch readout covering adoption, quality, issues, and next steps.
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?To handle and protect user data with storage, privacy, and compliance in mind, only the necessary user-provided inspiration, preferences, and interaction logs would be stored. Also, consent, encryption, retention limits, access controls, deletion options, and clear privacy disclosures would be implemented / applied. Legal would review privacy language, data handling, Instagram/manual upload flows, AI disclaimers, and travel advice boundaries. Audit logs would track prompt version, model version, output, and issue flags.
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?Current processes N/A for MVP - Developing a new product from 0 to 1 Before releasing, I would ensure that moderation, legal, and audit processes are in place. For the current MVP, moderation rules have been included to prevent unsafe, sensitive, or unsupported outputs, including fake bookings, invented prices, sensitive inference, unsafe travel certainty, and unsupported legal/visa guidance. Also, assuming this will be released in US and Canada first, I would also ensure the MVP complies with relevant state privacy laws (US) / PIPEDA (Canada), FTC Act (US) / Competiion Act (Canada), and ADA (US) / AODA (Canada), for example.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?User metrics: Activation rate, upload/connect completion, Travel Taste Profile completion, destination selection rate, itinerary generation rate, refinement rate, save/share rate, user satisfaction, repeat usage Business metrics: Itinerary-to-booking funnel performance, reduced planning time, user retention, and premium feature interest
AI MetricsHow will you measure AI performance and accuracy?In addition to the Golden Test Set score, I would measure AI performance and accuracy with hallucination rate, personalization score, itinerary feasibility score, constraint adherence, tone score, and human review pass rate.
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?Users will be able to access in-product help, FAQ, feedback form, and support email/chat. Escalation ownership will be clearly implemented where support handles user issues, product triages bugs, legal reviews privacy/safety issues, and engineering owns outages and system performance (e.g., latency).
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?Feedback will triaged and implemented by severity from P0-P4 and gathered across the following sources: critical safety/privacy issues, hallucinations, broken flows, poor recommendations, UX friction, and feature requests. Critical issues (i.e., P4) trigger immediate review and possible rollback within a day.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?Post-launch monitoring will include monitoring latency, failure rates, upload errors, hallucination flags, unsafe output flags, cost spikes / fluctuations, user drop-off, negative feedback, and support ticket themes to spot operational / AI issues after the launch.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?The team will review launch metrics on a weekly basis, re-run Golden Test Set after prompt/model/data changes, add newly identified real-world failures to the test set, update guardrails, and refresh the static knowledge pack as needed.
Download the .xlsx ↓