← All capstone projects

Local Commerce

Smart Offer Badges

Built by Ali Khosrojerdi Cohort 9 Local commerce / map-based offer discovery

Smart Offer Badges is an AI-powered local offer discovery and ranking concept layered on top of map-based business search. The project tackles the gap between seeing nearby businesses and understanding which real-time promotions are actually worth attention. Based on the visible slides and prototype, the system uses user query understanding, intent classification, AI evaluation, and multi-factor ranking across relevance, distance, and attractiveness, then presents results in a map and card interface with explanation cues.

The problem

Map-based local search shows nearby businesses but not which real-time promotions are actually worth attention. Users make high-intent decisions without incentives, and merchants have no way to influence discovery at the moment that matters. Once offers do exist, the hard problem is deciding which to surface without creating promotional clutter or irrelevant noise — poor offer selection and irrelevant promotions are the most frequent, most severe pain points. Traditional ad systems — Google Maps Local Ads, Apple Maps promotions, Yelp Ads — prioritize exposure and spend over relevance, against a backdrop of declining trust in digital promotions, signal overload, and privacy constraints.

The solution

Smart Offer Badges is an AI decision-and-ranking engine that layers onto map-based search to surface only offers genuinely worth showing. It interprets real-time intent and context, scores every nearby offer, suppresses low-quality ones, and ranks the rest across relevance, engagement, monetization, and exploration — with a diversity guard to prevent any one category from dominating. The engine is trust-first and conservative by default: when in doubt, it suppresses. It optimizes across relevance, engagement, and monetization simultaneously rather than maximizing exposure, and improves through a closed feedback loop that learns from impressions, clicks, and navigation. Results appear as ranked map badges with lightweight explanation cues.

How it works

The architecture is a hybrid: a constrained, instruction-following LLM handles intent interpretation, context understanding, and tradeoff reasoning under ambiguity, while a deterministic rule layer enforces eligibility, constraints, consistency, and auditability. The system prompt casts the model as a neutral platform arbiter that prioritizes user trust over merchant exposure, never optimizes for impressions or spend, and returns strict JSON only — an approved_offers array with confidence scores, no free-text reasoning. Inputs are structured, non-PII fields — offer metadata, merchant trust score, distance, inferred search intent, time context, policy flags — with low temperature for determinism and suppression by default when inputs are missing or conflicting. Manual evaluation across representative scenarios reached an ~85–90% pass rate, near 100% on high-confidence cases; early failures on ambiguous, borderline-trust queries were resolved by tightening suppression-first language and clarifying that trust outranks discount magnitude.

Who it's for

The product is B2B, sold to map-based search and local discovery platform operators (such as Google Maps and Apple Maps) who license the decision engine as a backend microservice. It integrates within existing map UI surfaces rather than introducing new navigation flows. It indirectly serves two end-user groups: consumers making immediate high-intent decisions, who benefit from relevance and reduced clutter, and local merchants creating offers, who benefit from qualified, high-intent visibility. Revenue is transactional and outcome-based — the platform pays per qualified high-intent action such as directions or visit intent, rather than for impressions or guaranteed ranking.

Why it matters

The target market is projected to grow at roughly 10–15% CAGR over 3–5 years, driven by MarTech (~15–20% CAGR) and location-based advertising, rising high-intent local search, and advances in real-time AI decisioning. The differentiator is a trust-first, self-optimizing marketplace that balances relevance and monetization while protecting the user experience. Rollout is staged and metrics-gated: a controlled pilot on select food and coffee merchants, then a live A/B test on ~5–10% of users comparing AI badges against a no-badge or rules-based baseline, then incremental scale (25% → 50% → 100%) only after meeting thresholds on relevance, latency, and high-intent actions — with instant feature-flag rollback to a safe "no-badge" mode if any metric degrades.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:[Ali Khoserojerdi]
Your Product:[Smart Offer Badges for Local Search ]
Your Industry:[MarkTech]
Date:[April,27,2026]
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Marktech Local advertising / AI‑driven promotionsPlease leave this area blank. This space is for the Instructor to provide you with feedback.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?Headwinds: declining trust in digital promotions, signal overload, privacy constraints, and monetization‑vs‑relevance tension. Tailwinds: growth of high‑intent local search, merchant demand for outcome‑based ads, and advances in real‑time AI decisioning. Competitors: Google Maps Local Ads, Apple Maps promotions, Yelp Ads, and other platform‑level local advertising solutions.
What is the projected growth rate of your target market segment over the next 3-5 years?The target market is projected to grow at approximately 10–15% CAGR over the next 3–5 years, driven by strong growth in MarTech (≈15–20% CAGR) and location‑based/local digital advertising (≈10–15% CAGR), supported by rising high‑intent local search and AI‑driven personalization.
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Startup (early‑stage / concept‑to‑MVP)
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)The business monetizes via a transactional, outcome‑based model, selling an AI decision engine to map‑based search platforms that evaluate merchant participation and generate revenue per qualified high‑intent action (e.g., directions or visit intent), rather than guaranteed ranking or impressions.
Who is your primary customer base (B2B, B2C, B2B2C)?B2B — map‑based search and local discovery platform operators.
DifferentiatorsWhat are the key differentiators for your company?Differentiated by an AI-powered marketplace ranking engine that optimizes across relevance, engagement, and monetization simultaneously. Unlike traditional ad systems that prioritize exposure or spend, this system: - selectively surfaces high-value offers - learns continuously from user behavior - balances fairness through diversity constraints - adapts ranking dynamically in real time This creates a trust-first, self-optimizing local search marketplace.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?Map‑based search platform operators (e.g., Google Maps, Apple Maps) who license an AI decision engine to intelligently surface merchant offers in local search.
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?End users include: (1) consumers using map‑based local search to make immediate decisions, and (2) local merchants creating promotional offers. Primary revenue‑impacting users: map‑based search platform operators who license the AI decision engine. Goals: consumers seek relevance and trust in high‑intent moments; merchants want qualified outcomes (visits/actions); platforms aim to monetize local intent without degrading user experience.
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?Core features include: an AI-powered marketplace ranking engine that evaluates, filters, and ranks merchant offers in local search using multi-factor scoring. The system combines: - AI relevance scoring (intent + context) - Dynamic ranking (quality + bid + exploration) - Engagement-driven learning loop (impressions, clicks, navigation) - Revenue optimization using expected revenue (bid × adjusted CTR) - Session-based personalization (recent queries + category preference) - Diversity guard to prevent category dominance - Exploration layer to surface underexposed offers This enables platforms to move beyond static
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)Primarily for map‑based search platforms (e.g., Google Maps, Apple Maps) that use the AI decision engine to optimize offers ranking in local search; it indirectly serves end users (consumers) by improving relevance and merchants by delivering qualified, high‑intent visibility.
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?Map‑based search platforms integrate the AI decision engine, configure relevance and monetization guardrails, and enable merchant offers; as users search locally, the AI evaluates real‑time intent and context to selectively surface the most relevant offers, driving higher engagement and monetization while preserving user trust and reducing promotional clutter.
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?Major unmet needs include the lack of native promotional offers in map‑based search, forcing users to choose without incentives and merchants without a way to influence high‑intent discovery; once offers exist, key friction points are deciding which offer to surface, avoiding promotional clutter, and ensuring relevance—making poor offer selection and irrelevant promotions the most frequent and severe pain points.
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.Most severe pain points include selecting the most relevant offers at moments of intent, avoiding promotional clutter, explaining why an offer was shown, handling cold start offers, and reducing reliance on manual rules—each requiring LLM powered reasoning under ambiguity across context interpretation where signals conflict, intent, and ambiguity. Requires AI reasoning to resolve ambiguous intent and conflicting contextual signals where deterministic rules are insufficient. Additional opportunity: moving beyond binary offer selection into a continuous ranking and optimization system that learns from user interactions, balances monetization with relevance, and adapts dynamically over time.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.AI-powered local search marketplace engine that evaluates, ranks, and dynamically surfaces merchant offers based on real-time context, user intent, and engagement signals. The system optimizes across multiple objectives: relevance (AI score) user engagement (CTR, navigation signals) monetization (bid × adjusted CTR) fairness (category diversity constraints) exploration (surfacing underexposed offers) The engine evolves continuously through a closed feedback loop, incorporating impressions, clicks, and navigation behavior to refine confidence scores, personalize results, and improve ranking quality over time.
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.Ranked by impact and feasibility: (1) AI‑driven multi‑offers selection using real‑time intent and context inference (highest impact, highest feasibility); (2) semantic offers understanding and normalization to enable relevance and cold‑start scoring; (3) natural‑language explanation generation to build trust and adoption. Primary focus for this project: the AI decision engine that selects and surfaces the single most relevant offer per search moment.
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?Map-based search platforms enable merchant offers and configure AI guardrails. When a user performs a search: 1. The AI interprets real-time intent and contextual signals 2. All nearby offers are evaluated and scored 3. Low-quality offers are suppressed 4. Approved offers enter a ranking engine Ranking is computed using: - relevance (AI score) - engagement signals (confidence + CTR) - monetization (bid) - exploration (uncertainty boost) A diversity guard ensures category balance before final results are shown. Users interact with offers (click or navigation), generating signals that update: - confidence scores - user preferences - CTR estimates The system continuously re-ranks and improves via a closed feedback loop while preserving user trust and optimizing for high-intent actions.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?The AI operates primarily as a backend decision and ranking engine integrated within existing map-based UI surfaces. It consumes contextual inputs (query, location, time, movement) and merchant offers, then outputs ranked results rather than a single decision. The UI displays: - ranked offers as map badges - visual prioritization (position reflects ranking) - lightweight explanation signals (e.g., why ranked #1) No new navigation flows are introduced; instead, the AI enhances existing search results with dynamic ranking, personalization, and monetization-aware ordering.
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?The prototype demonstrates the end-to-end AI ranking system: Inputs: - user query, time, location, movement - merchant offers Processing: - intent classification - AI relevance scoring - suppression filtering - ranking engine (quality + bid + exploration) - diversity guard - revenue signal computation Outputs: - ranked list of offers - approved vs suppressed offers - score breakdown (AI, preference, confidence, bid, exploration) - explanation of why top results are shown Essential launch features: - intent/context scoring - suppression filtering - ranking engine Deferred: - deeper personalization models - advanced explainability layers
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?System Role You are an AI decision engine embedded inside a map‑based local search experience. Primary Responsibility Your sole responsibility is to determine which merchant promotions, if any, deserve to appear as contextual offer badges within search results. Guiding Principles You must prioritize user trust and contextual relevance above merchant exposure at all times. You do not optimize for impressions, clicks, rankings, or advertiser spend. You must behave conservatively by default. When in doubt, choose suppression. Decision Logic You must evaluate each available merchant offer independently and approve only those that meet a high confidence threshold for contextual relevance based on the user’s inferred intent, location, timing, and situational signals. It is acceptable to approve multiple offers only if each approved offer is clearly beneficial to the user in the current search context. You must suppress all offers that do not meet this standard. Assumptions & Constraints All signals are inferred probabilistically from available contextual information. You must not assume explicit user preferences, personal attributes, or long‑term behavioral profiling. If no promotion is clearly beneficial to the user at this moment, you must return no approved offers. You do not explain your reasoning to the user. Input You will receive structured input that may include: A search query Inferred contextual signals (for example: location, time, movement, environment) A bounded list of available merchant offers with defined attributes and constraints Output (Strict JSON Only) The output must be a deterministic decision object in the following shape: approved_offers: an array of approved offers Each approved offer contains: offer_id (string) confidence_score (number between 0.0 and 1.0) Rules Always return valid JSON. Never include additional fields or explanatory text. The approved_offers array may be empty. Only include offers that independently meet the relevance threshold. Never invent, modify, or infer offer attributes beyond what is explicitly provided.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?Because the AI functions as a decision engine rather than a generative assistant, output quality is evaluated based on decision correctness, consistency, and trust preservation rather than linguistic creativity. Relevance: Approved offers must clearly align with the user’s inferred search intent and situational context such as location, timing, and environment. Each approved offer must independently justify its inclusion, while weak, generic, or marginally related offers should be suppressed. Accuracy and Eligibility: All approved offers must strictly respect provided constraints including validity windows, locations, and usage conditions. The system must never approve expired, invalid, or inapplicable offers, and must not invent, alter, or infer offer attributes beyond what is explicitly provided. Conservatism and Hallucination Avoidance: The system must default to suppression when relevance or confidence is insufficient. It must not infer personal user preferences, demographic traits, or historical behavior, and must never fabricate offers, discounts, or merchant details. Confidence Calibration: Confidence scores must meaningfully reflect the strength of contextual relevance. High confidence approvals should be deliberate and rare, while multiple offers may be approved only when each independently meets the relevance threshold. Output Consistency and Determinism: Outputs must always conform to the defined JSON schema. Identical inputs should yield identical outputs, and the approved_offers array must be empty when no offer qualifies. Tone and Presentation: The system does not generate user-facing language or explanations. All outputs remain neutral, non-editorial, and purely decision-oriented. User Trust Preservation: Promotional content should feel helpful and contextual rather than ad-like. The system prioritizes fewer, higher-quality offers over maximum exposure. A “good” output is a deterministic decision object that includes only clearly relevant and eligible offers or none at all, reflects appropriate confidence, obeys all constraints, and preserves user trust through conservative approval and strong suppression behavior.
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?Test prompts should cover a representative set of real‑world scenarios to validate that the AI decision engine consistently prioritizes relevance, accuracy, and user trust while suppressing inappropriate promotions. Core use cases include high‑intent searches with strong contextual alignment, such as a morning “coffee near me” query during commuting hours with valid breakfast‑time café offers, or a “lunch nearby” search around noon with multiple restaurants offering time‑appropriate discounts. These cases validate that the system correctly approves one or more clearly beneficial offers when intent, timing, and offer constraints align. Additional standard use cases include situational relevance, such as bad‑weather conditions increasing the relevance of nearby indoor venues, late‑night searches aligning with extended‑hour food offers, or proximity‑based searches where multiple nearby merchants independently meet relevance thresholds. In these cases, the system should approve multiple offers only when each independently qualifies. Edge cases should test ambiguity and boundary conditions. Examples include searches with vague or exploratory intent, such as “things to do nearby,” where offers exist but lack clear relevance, or scenarios where offers are only marginally applicable due to timing, distance, or weak intent signals. Edge cases also include situations where multiple offers are theoretically relevant but differ in strength, requiring the system to suppress borderline options rather than lower the bar for inclusion. Other edge cases include conflicting contextual signals, such as a user searching while in transit, unclear movement state, rapidly changing location, or overlapping offers with similar content but different constraints. The system should remain conservative and avoid over‑approval in these scenarios. Negative cases are critical and must explicitly validate suppression behavior. These include expired offers, offers outside valid time windows, offers that violate location or usage constraints, and promotions unrelated to the search query despite proximity. The system should also suppress offers when confidence scores are low, when context is insufficient to infer intent, or when showing a promotion could confuse or mislead the user. Additional negative cases include attempts to surface merchant promotions that rely solely on availability or advertiser presence rather than relevance, and any scenario in which approving an offer would prioritize exposure over user benefit. Together, these test cases ensure the system correctly approves relevant promotions, suppresses invalid or weak offers, handles ambiguity conservatively, and consistently defaults to protecting user trust over promotional volume.
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?This solution is best served by a constrained, instruction‑following large language model (LLM) used as a decision engine rather than a generative system. The model’s role is to evaluate multiple contextual signals—offer quality and freshness, merchant trust indicators, spatial and temporal context, and inferred user intent—and decide whether an offer badge should be shown without degrading relevance or user trust. Its key capabilities include context-aware reasoning across heterogeneous inputs, effective performance in low‑data or zero‑shot scenarios, and the ability to produce strictly structured, deterministic JSON outputs that integrate cleanly with map platforms. Limitations include potential non-determinism, latency, and cost, which are mitigated through low-temperature settings, strict output schemas, input validation, and execution only on high‑intent events. The model integrates as a backend microservice that receives structured inputs from the map platform and returns an explainable show/suppress decision with confidence and reason codes, without owning UI or rendering logic. LLM used for: intent interpretation The system combines deterministic eligibility rules with LLM-based contextual reasoning, where the LLM interprets intent and resolves ambiguity while the rule layer enforces consistency, safety, and auditability. context understanding tradeoff reasoning deterministic layer used for: constraints eligibility enforcement The system extends beyond decision-making into a hybrid ranking framework, where LLM-based reasoning is combined with deterministic ranking logic and marketplace dynamics. This includes integration of learning signals, bid-based monetization, exploration strategies, and feedback-driven adaptation, evolving the system into a real-time optimization engine rather than a static decision model.LLM used for: intent interpretation The system combines deterministic eligibility rules with LLM-based contextual reasoning, where the LLM interprets intent and resolves ambiguity while the rule layer enforces consistency, safety, and auditability. context understanding tradeoff reasoning deterministic layer used for: constraints eligibility enforcementPlease leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.The AI requires structured, non‑UI input fields including: business_id (string, platform internal ID, required), offer_text (string, merchant‑provided promotion description, required), offer_type (enum e.g., discount/freebie/BOGO, merchant CMS or ad system, required), offer_expiry (ISO‑8601 date, merchant CMS, required), merchant_trust_score (float 0–1 or tiered enum, platform trust system, required), merchant_category (string/category ID, platform taxonomy, required), distance_to_user (numeric in meters, map engine, required), user_search_intent (string or intent label inferred from query, search engine, required), time_context (timestamp + daypart label, system clock, required), historical_offer_performance (aggregated metrics object, analytics system, optional), and platform_policy_flags (boolean or enum flags, policy service, required). All inputs are machine‑readable, privacy‑safe, and passed in a strict JSON schema to enable deterministic decision output.
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?Yes. Optional and platform‑configurable fields include relevance_threshold (numeric confidence cutoff defined by the map platform), suppression_bias (enum favoring trust‑first vs. revenue‑balanced behavior), max_badge_density (numeric cap per viewport), preferred_offer_types (ranked list set by platform or user), and historical_user_interaction_signals (aggregated clicks, dismissals, or conversions when available). These fields do not override core relevance logic but act as weighting modifiers that influence decision confidence, suppression strictness, and prioritization when multiple eligible offers compete. If absent, the AI defaults to a conservative trust‑preserving baseline to ensure consistent and safe outputs.
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)Output quality is judged by whether the AI produces a strictly valid JSON response with all required fields populated, a clear and consistent show/suppress decision, and an explainable reason code aligned with platform policy. The decision must be contextually relevant to the user’s search intent and location, factually grounded in provided input signals, conservative in ambiguous cases to protect user trust, and consistent across similar inputs under the same configuration. “Good” output prioritizes relevance and credibility over exposure, avoids unnecessary badge clutter, and demonstrates stable behavior that can be audited, measured, and tuned by the platform.
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?Yes. Optional and platform‑configurable fields include relevance_threshold (numeric confidence cutoff defined by the map platform), suppression_bias (enum favoring trust‑first vs. revenue‑balanced behavior), max_badge_density (numeric cap per viewport), preferred_offer_types (ranked list set by platform or user), and historical_user_interaction_signals (aggregated clicks, dismissals, or conversions when available). These fields do not override core relevance logic but act as weighting modifiers that influence decision confidence, suppression strictness, and prioritization when multiple eligible offers compete. If absent, the AI defaults to a conservative trust‑preserving baseline to ensure consistent and safe outputs. Additional optional signals include: - engagement metrics (clicks, navigation) - session-level query history - inferred category preferences These inputs enable personalization, confidence adjustment, and ranking optimization without relying on persistent user identity.
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.The starting system prompt instructs the model to act as a conservative, trust‑first offer eligibility decision engine for map‑based search, explicitly prohibiting UI generation, persuasive language, or merchant favoritism and requiring deterministic, policy‑aligned judgment. Initial instructions define the persona as a neutral platform arbiter optimizing user relevance and trust over revenue, with inputs limited to structured JSON fields such as offer metadata, merchant trust signals, spatial context, and inferred search intent, and outputs constrained to a strict JSON schema containing show/suppress decisions, confidence scores, and standardized reason codes. Constraints include low temperature, no free‑text explanations, no use of external knowledge, and suppression by default in ambiguous cases. Prompt variations to be tested include stricter vs. relaxed suppression bias, different confidence threshold definitions, alternate reason‑code granularity, and reordered input emphasis (e.g., trust‑first vs. intent‑first). Performance optimization techniques include prompt tightening to reduce ambiguity, schema validation with automatic retries, rejection sampling for malformed outputs, offline evaluation against human‑reviewed edge cases, and continuous tuning of weights using aggregated interaction feedback.
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?Yes, the prompt is revised iteratively based on evaluation results and edge‑case review. Revisions typically include clarifying suppression‑first language to reduce false positives, tightening definitions of trust and relevance signals, refining reason‑code mappings for better auditability, and strengthening constraints on determinism and schema adherence. Each change is made to improve consistency, explainability, and alignment with platform trust goals rather than output creativity. Prompt evolution is tracked through versioned system prompts stored alongside model configuration metadata, with changes logged using structured changelogs that capture the rationale, expected behavioral impact, and evaluation results. Performance is monitored using before‑and‑after comparisons on a fixed test set of human‑reviewed scenarios and key metrics such as suppression accuracy, consistency across similar inputs, and override rates.
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?The system relies on multiple structured data sources provided by map platforms and merchants, including offer metadata from merchant CMS or ad systems, merchant trust and policy signals from platform risk services, spatial and temporal context from the map engine, inferred search intent from query processing, and aggregated historical interaction data from analytics pipelines. No raw PII or free‑text user data is used. Data preparation focuses on normalization and structuring rather than model training: inputs are cleaned for missing or stale fields, validated against strict schemas, standardized into enums and numeric ranges, and enriched with derived signals such as distance bands or intent labels. Evaluation data consists of logged decision scenarios paired with human‑reviewed judgments and downstream user outcomes (clicks, dismissals, navigation). This system does not rely on RAG; however, if introduced, retrieval would be limited to small, policy‑controlled reference documents such as platform rules or offer eligibility guidelines, chunked by rule section, embedded using sentence‑level embeddings, and retrieved deterministically to inform reasoning without exposing the model to external or dynamic content. Inputs are interpreted both deterministically and via LLM reasoning to capture semantic meaning and contextual nuance.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.The most common inputs are structured fields passed as JSON, including business_id (“cafe_48219”), merchant_category (“Coffee Shop”), offer_text (“10% off any drink”), offer_type (“percentage_discount”), offer_expiry (“2026‑05‑15”), merchant_trust_score (0.92), distance_to_user (240 meters), user_search_intent (“coffee near me”), time_context (“weekday_morning”), and platform_policy_flags (“eligible”). Based on these inputs, the expected output is a deterministic JSON decision such as show_offer:true with confidence_score:0.81 and reason_code:“HIGH_INTENT_RELEVANT_TRUSTED_OFFER”. In contrast, a low‑quality or misaligned input (e.g., expired offer, trust_score 0.45, distance 1.8 km, generic intent “restaurants”) would typically produce show_offer:false with a lower confidence score and a suppression reason such as “LOW_RELEVANCE_OR_TRUST”. Outputs are always machine‑readable, policy‑aligned, and explainable, enabling consistent rendering or suppression by the map platform.
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)Boundary‑testing examples include scenarios with missing or degraded inputs such as absent trust scores or incomplete offer metadata, ambiguous intent cases like a generic query (“food near me”) combined with multiple nearby offers, and conflicting signals such as a high discount from a low‑trust merchant or a trusted merchant with an almost‑expired offer. Out‑of‑domain cases include offers that fall outside supported categories, non‑local promotions, or promotional language that violates platform policy. The AI is expected to default to suppression in all such cases, returning clear suppression reason codes rather than attempting inference, thereby validating safe failure behavior, conservative defaults, and robustness under uncertainty.
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?Sample inputs were run through the system prompt using representative local search scenarios and reviewed manually against relevance, trust protection, consistency, and explanation criteria. In most standard cases (high‑trust merchant, fresh offer, strong local intent), outputs met expectations by correctly showing the offer with consistent confidence scores and appropriate reason codes. Failures primarily occurred in edge cases with ambiguous or incomplete inputs, such as generic search intent combined with borderline trust scores, where early prompt versions occasionally surfaced offers that reviewers felt should be suppressed. These failures were traced to insufficiently strict suppression language and unclear weighting between intent and trust signals. Revisions tightened default suppression behavior and clarified priority ordering, after which outputs aligned with manual expectations and consistently favored trust‑preserving decisions in uncertain scenarios.
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?In manual evaluation across a representative test set of core scenarios, the AI achieved an overall pass rate of approximately 85–90% against primary criteria including relevance, trust preservation, consistency, and explainability. High‑confidence cases (strong local intent, high‑trust merchants, valid offers) passed at near 100%, while failures were concentrated in ambiguous edge cases such as generic queries with borderline trust scores or competing signals, where initial prompt versions surfaced offers that reviewers expected to be suppressed. After prompt tightening, suppression accuracy in these edge cases improved significantly, with remaining failures deemed acceptable conservative tradeoffs rather than harmful errors.
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?Edge cases identified during testing include missing or degraded inputs such as absent merchant trust scores or incomplete offer metadata, ambiguous user intent (e.g., generic queries like “restaurants near me”), and conflicting signals where high‑value offers originate from low‑trust merchants or trusted merchants present weak or near‑expired promotions. Additional edge cases include high badge density in dense urban areas, long distance despite strong intent, category mismatches between search intent and merchant type, and out‑of‑domain promotions that violate platform policy or fall outside supported verticals. In all such cases, the system is designed to fail conservatively by suppressing offers and returning clear suppression reason codes to preserve user trust and predictable behavior.
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?Based on failures, feedback, and edge‑case observations, several system prompt and configuration adjustments were made, including strengthening suppression‑first language to reduce false positives in ambiguous scenarios, clarifying the priority ordering of trust signals over discount magnitude and merchant spend, and tightening definitions of relevance for generic or out‑of‑domain search intents. Additional constraints were added to enforce stricter schema adherence, require explicit reason codes for all suppress decisions, and default to suppression when required inputs are missing or conflicting. Confidence thresholds were recalibrated to better handle borderline trust scores, and input weighting was adjusted to reduce sensitivity to promotional aggressiveness. These changes improved consistency, auditability, and alignment with the platform’s trust‑preserving goals while maintaining acceptable performance in high‑confidence cases.
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?The primary evaluation approach is human review, used to assess qualitative criteria such as trust preservation, relevance, and explanation quality that are not reliably captured by automated metrics in early‑stage AI products. Human reviewers evaluate outputs against predefined pass/fail criteria using representative scenario sets. To scale testing, human‑labeled baseline datasets are combined with scripted evaluations that check schema validity, decision consistency, confidence thresholds, and suppression defaults across large synthetic and replayed input sets. As coverage grows, selective model‑grader checks and batch simulations are used to stress‑test edge conditions, while human review remains focused on ambiguous or high‑risk cases, enabling scalable validation without sacrificing judgment quality.
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?Evaluations are re‑run on a regular cadence tied to meaningful changes in the system: immediately after any prompt or system instruction update, periodically when new data distributions or merchant behaviors emerge, and continuously post‑launch through monitoring of logged decisions and outcomes. Lightweight automated checks (schema validity, consistency, threshold adherence) run on every deployment and batch simulation, while human reviews are conducted on a scheduled basis (e.g., weekly or bi‑weekly) for new edge cases and high‑risk scenarios. After launch, evaluations are triggered proactively by drift indicators such as changes in suppression rates, override frequency, or user interaction patterns, ensuring the system remains aligned with trust and relevance goals over time.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?Yes — infra is fully tested and documented for MVP readiness: APIs are validated via schema and edge-case testing (including malformed inputs and retry handling), rate limits and throttling are enforced and tested under simulated load, monitoring is implemented with alerts on latency/error thresholds, and rollback is enabled via prompt versioning + feature flags with successful fail-safe fallback to “no-badge” mode; all components (API contracts, decision logic, failure modes, and escalation paths) are documented and accessible to engineering, product, and support teams.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?Yes — internal teams are trained and documentation is complete at MVP level: support, comms, and legal have been onboarded through structured walkthroughs covering system behavior, guardrails, escalation procedures, and policy boundaries, with scenario-based examples (approved vs suppressed offers); comprehensive documentation is finalized, including API contracts, decision logic principles, allowed/disallowed monetization practices, failure modes, and support playbooks, with a shared knowledge base ensuring all teams have consistent, up-to-date reference material.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?Staged rollout: controlled pilot → A/B test → gradual scale — initial access is limited to a small, geo/category-constrained pilot (e.g., select food/coffee merchants) using low-risk traffic or synthetic data; once stable, a live A/B test is launched on ~5–10% of real users comparing AI-driven badges vs baseline (no-badge or rules-based), with exposure gated by high-intent queries only; broader access is granted incrementally (25% → 50% → 100%) only after meeting predefined success thresholds on relevance, latency, and high-intent actions, with the ability to instantly rollback via feature flags if performance degrades.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?Scale readiness is ensured through controlled capacity gating, stateless architecture, and real‑time performance thresholds: the AI decision service is horizontally scalable (stateless APIs + cached context patterns), with predefined SLAs for latency, error rate, and output validity; readiness is validated via load simulation testing and enforced through strict rate limits and query gating (initially high-intent searches only), ensuring the system can handle incremental traffic without degradation. Monitoring and scale-up are metrics-driven and staged: initial volume is tightly capped and tracked using real-time dashboards (latency, error rates, suppression ratio, approval distribution, and high-intent actions), with alerting on anomalies; scale-up follows a gated ramp (e.g., 5% → 25% → 50% → 100%) only when stability thresholds are consistently met (e.g., <1% error rate, stable latency under SLA, high relevance scores), with the ability to pause or rollback instantly via feature flags if any metric deviates.
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?External assets are prepared as a focused, trust-first communication package: includes a concise product one-pager explaining how AI selects relevant offer badges and monetization logic, a clickable demo prototype (Google Maps–style) showcasing real user scenarios and decision outcomes, a short explainer video highlighting value for users and merchants, and an FAQ addressing transparency (why some offers are shown/suppressed), pricing model (pay-per-high-intent action), and policy safeguards; additionally, developer/API guides and partner onboarding docs are created for platform integration, ensuring external stakeholders clearly understand functionality, guardrails, and integration steps.
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?share a simple launch brief upfront (what we’re launching, who sees it, and success criteria), then provide quick weekly updates with key metrics and learnings during pilot and A/B test; use a shared dashboard or document for visibility into performance (e.g., engagement, suppression rate, errors), and align quickly with engineering if issues come up—no heavy process, just fast feedback loops and clear visibility.
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?user data is minimized by design — the system only processes non-PII inputs (query context, location signals, and offer metadata) and does not store or rely on personal user information; any logged data is anonymized and used only for debugging and improving decision quality, with strict access control; no long-term user profiling is done, and all data handling aligns with platform privacy standards and basic compliance practices (no sensitive data collection, clear audit trail, and safe fallback if anything is uncertain).
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?content moderation is handled through built-in guardrails that filter out misleading, expired, or irrelevant offers at decision time; legal alignment is covered by limiting to non-PII data and following standard advertising and platform policies (no deceptive or unsafe promotions); and auditability is ensured through simple logging of inputs, outputs, and decision reasons so any issue can be traced and reviewed — overall compliant for MVP scope, with room to formalize further as the product scales.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?Success is measured by: User metrics: - adjusted CTR (position bias corrected) - navigation rate (high-intent actions) - reduced badge clutter (suppression effectiveness) Business metrics: - revenue per impression - revenue per session - high-intent action rate (e.g., directions) Marketplace metrics: - diversity balance across categories - exploration success (performance of new/low-data offers) - ranking quality and stability over time
AI MetricsHow will you measure AI performance and accuracy?AI performance is measured across: - Relevance accuracy (alignment with intent) - Suppression precision (avoiding low-quality offers) - Adjusted CTR (corrected for position bias) - Confidence calibration (alignment between prediction and outcome) - Learning effectiveness (improvement over time) - Exploration efficiency (performance of underexposed offers) - Ranking consistency across similar inputs These metrics ensure the system continuously improves while balancing relevance, monetization, and fairness.
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?users and partners are supported through the platform’s existing support channels, with clear ownership assigned to product + engineering for issues related to AI decisions; every decision is logged, so support can quickly trace what happened, and there’s a defined escalation path (support → product → engineering) to resolve issues fast without ambiguity.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?when businesses raise complaints, we rely on logged decision data to explain why their offer wasn’t shown (e.g., low relevance, weak offer value, or policy constraints); instead of manual overrides, we provide clear guidance on how to improve eligibility (better targeting, stronger offers, clearer metadata); repeated patterns in complaints are treated as signals to refine the model or rules, but we protect user trust first — no guarantee of visibility, only fair and consistent evaluation.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?every badge decision (show/suppress) is logged with query context (e.g., “coffee near me”), merchant offer data, confidence score, and reason (e.g., relevance, value, distance), allowing full traceability of why a business was or wasn’t shown; dashboards track not just technical metrics (latency, errors) but core product signals like badge CTR, high-intent actions (directions), suppression rate (avoiding spam), and merchant distribution to ensure fairness; alerts are triggered on issues like too many low-quality offers being shown, overly aggressive suppression (good offers missing), or inconsistent decisions across similar queries; this ensures we’re not just monitoring system health, but actively protecting user trust, offer relevance, and marketplace balance in real time.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?Continuously collect data from badge performance (CTR, directions), suppression patterns, merchant complaints, and logged decisions, then review it in lightweight weekly check-ins to spot trends (e.g., good offers being missed or weak ones getting through); issues are categorized (model logic, data quality, policy gaps) and prioritized based on user trust impact; updates are rolled out through prompt tuning, threshold adjustments, or rule refinements—shipped quickly via versioning and feature flags—so the system keeps improving without heavy process or downtime.
Download the .xlsx ↓