← All capstone projects

SaaS

Pulse AI

Built by Vijeta Marwah Cohort 9 B2B martech and customer lifecycle optimization

Pulse AI is a decision intelligence layer for lifecycle and growth product managers at high-volume consumer businesses. It combines LLM-generated semantic understanding of customer events, CRM context, and voice-of-customer logs with machine learning to recommend send or suppress decisions, channels, timing, and suggested copy for individual customers. The product also exposes confidence scores, explanations, and a dashboard for incremental value measurement.

The problem

Enterprises spend billions on customer engagement across push, email, SMS, WhatsApp, in-app messaging, and loyalty programs — but most systems optimize campaigns, not customer outcomes. Platforms like Braze, MoEngage, Salesforce, and WebEngage are rule-based orchestration tools: they send, but they don't decide whether a message should be sent at all. The result is wasted spend, communication fatigue, opt-outs, and margin erosion. Lifecycle PMs can't tell which message drives opt-outs, whether suppressing a comm loses revenue, or that a customer with a poorly settled claim is being blasted with marketing. As one PM put it: "I'm rebuilding the same lifecycle in Braze every quarter, and when leadership asks if our 9am send drove revenue I'm guessing."

The solution

Pulse AI (Pulse Customer Intelligence) is a customer behavioral intelligence layer that sits in the white space between orchestration and customer-level decisioning. Rather than optimizing campaigns, it optimizes customer outcomes — recommending the highest-value intervention for each individual user. For every user, Pulse recommends send, suppress, wait, best channel, or incentive-vs-no-incentive, backed by suppression intelligence and behavioral state scores for purchase intent, churn risk, and fatigue. Each recommendation carries a predicted impact score, confidence level, expected incremental value, and a plain-language reasoning summary. The MVP focuses on intelligence, not automation: a human stays in the loop to execute or reject.

How it works

Pulse builds behavioral intelligence across five layers — semantic event understanding, behavioral state intelligence, intervention prediction, suppression intelligence, and a continuous feedback loop. It is deliberately cost-aware: the runtime hierarchy is raw event → taxonomy lookup → embeddings/rules → LLM only when unresolved, so models are called for ambiguity rather than for every user. Models are assigned by task: gpt-5.1 for event taxonomy classification and Voice-of-Customer reasoning, gpt-5-mini for runtime fallback and NBA explanation, and gpt-4o-transcribe for audio. All LLM outputs are strict JSON with confidence scores. Customer protection rules (quiet hours, NPS thresholds, open support tickets) and business limits override any recommendation. A golden suite of 50 synthetic scenarios reached a 98% executable pass rate, with 100% hard-constraint compliance.

Who it's for

Pulse is B2B, targeting enterprises across ecommerce, fintech, travel, subscription, and food delivery with high communication volumes and existing CRM systems. Within the account, the economic buyer is the VP of Growth, the primary user is the Lifecycle PM, the technical buyer is data/martech, and legal/security/compliance approves. Primary users are CRM and lifecycle marketing leaders responsible for retention and engagement; their goals are more revenue through better targeting, reduced marketing spend, and higher customer LTV. Revenue is a subscription fee plus usage-based pricing per 1,000 MAUs scored.

Why it matters

The stakes are wasted marketing budget, customer churn from over-contact, and the inability to prove which sends actually drive incremental revenue. Pulse targets enterprise marketplaces growing at a projected 20–50% over five years, with no established competitor occupying the customer-level decision-intelligence layer. Rollout is deliberate. Each new client goes through an exhaustive data-connection setup and a 7-day observe period to solve the cold-start problem, then a narrow pilot (one journey, 5,000–50,000 users, a control-vs-test comparison) before 100% rollout. Pulse starts in recommendation mode; autonomous execution is unlocked only once enterprise confidence is established. The defensible moat is behavioral intelligence and longitudinal memory, not commoditized copy generation or orchestration.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Vijeta Marwah
Your Product:Pulse Customer Intelligence (Please see Executive Readout here)
Your Industry:Marketing Tech
Date:8th May 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Marketing Tech: B2B enterprise AI decisioning for lifecycle engagement The initial ICP includes ecommerce, fintech, travel, subscription, and food delivery companies with high communication volumes and existing CRM systems.Vijeta — the sharpest thing here is the category-level reframe from "campaign optimization" to "customer outcome optimization." That's not a feature pitch, it's an argument, and the 5-layer behavioral intelligence stack (semantic events → behavioral state → intervention prediction → suppression → feedback loop) actually backs it up with structure. That's the muscle this capstone is trying to build. Two pushes. 1. The persona is too abstract. "Growth Marketing Managers in large enterprises" is a job-title cluster, not a person. Who specifically — the CRM lead at a fintech whose push-open rates just dipped? The lifecycle PM at a food-delivery app already running Braze who can't tell which step in her onboarding flow is over-sending? Anchor one named archetype with a Tuesday-morning workflow; it sharpens every downstream design call. and 2, the pain points read like market analysis, not lived experience. "Cohort-based targeting" and "campaign-level optimization" are strategic gaps you've named from the outside. The actual pain is closer to "I'm rebuilding the same lifecycle in Braze every quarter, I can't tell which message is causing the opt-outs, and when leadership asks if our 9am send drove revenue I'm guessing." Re-write pain in the user's voice, not the analyst's.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?KEY PAIN POINT & SCOPE OF PROBLEM: Enterprises spend billions on customer engagement across push, email, SMS, WhatsApp, in-app messaging, and loyalty programs. Despite this, most systems optimize campaigns rather than customer outcomes.Platforms like Braze, MoEngage, Salesforce, and WebEngage primarily function as orchestration and campaign tools. They are rule-based, static, and campaign-centric. Marketing teams send massive volumes of messages daily, but a large portion is unnecessary. Many users would have converted organically or were over-contacted. This results in wasted spend, lower engagement, and margin erosion. Communication fatigue is also increasing, leading to opt-outs and churn.At the core, enterprises lack a unified behavioral understanding of customers across fragmented data systems. There is no established competitor in this space.Braze, MoEngage, Salesforce, and CDPs all adjacent. PULSE is a "white space between orchestration and customer-level decision intelligence Tailwinds: Channel costs such as WhatsApp/SMS costs, AI readiness, margin pressure.
What is the projected growth rate of your target market segment over the next 3-5 years?The initial ICP includes ecommerce, fintech, travel, subscription, and food delivery companies with high communication volumes and existing CRM systems. The projected growth rate for such enterprises marketplaces is 20-50% over the next 5 years.
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Pulse is a 0-to-1 MVP/pilot-stage product targeting scale-up and mature enterprises.
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)Subscription fee (to cover for fixed infra costs and integration efforts) and an ongoing usage based pricing based on MAUs Example: Subscription fee of $Y/month + $X per 1,000 MAUs scored
Who is your primary customer base (B2B, B2C, B2B2C)?B2B enterprises across ecommerce, fintech, travel, subscription, and food delivery companies with high communication volumes and existing CRM systems. Economic buyer = VP Growth; Primary User = Lifecycle PM; Technical buyer = data/martech Approver = legal/security/compliance
DifferentiatorsWhat are the key differentiators for your company?Pulse Intelligence is not a campaign or marketing automation tool. It is a customer behavioral intelligence layer. It continuously learns customer state, predicts responsiveness, and recommends the intervention with the highest long-term business value. The defensibility does not come from campaign optimization , messaging, orchestration, or AI copy generation. These are commoditized. The real moat is behavioral intelligence, intervention outcome learning, longitudinal memory, and feedback-driven optimization of customer value. Behavioral intelligence is built across five layers. 1. Semantic Event Understanding converts fragmented event data into unified behavioral signals. 2. Behavioral State Intelligence estimates intent, fatigue, and churn risk in real time. 3. Intervention Prediction evaluates the impact of different actions before execution. 4. Suppression Intelligence prevents harmful or low-value communication. 5. A continuous feedback loop improves all predictions over time.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?B2B enterprises across ecommerce, fintech, travel, subscription, and food delivery companies with high communication volumes and existing CRM systems. Economic buyer = VP Growth Primary User = Lifecycle PM
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?Persona 1: Secondary users include Lifecycle PMS driving revenue. Persona 2: Primary users are CRM and lifecycle marketing leaders responsible for retention and engagement. Their key goals are generating more revenenue (through improved conversion/CTR) , improve marketing efficiency (through reduced spends and better targeting) and increase customer LTV/retention through higher-quality engagement.
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?Pulse ingests behavioral data, normalizes event semantics, and builds a real-time behavioral model of each user. It predicts intent, fatigue, and responsiveness, then evaluates possible interventions to recommend the highest-value action for every customer.A feedback loop continuously improves prediction accuracy over time. The MVP focuses only on intelligence, not automation. It includes event ingestion, semantic normalization, behavioral state modeling, intervention recommendation, suppression intelligence, and a feedback learning loop. It also provides explainability for every recommendation.
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)Persona 1: Secondary users include Lifecycle PMS driving revenue. Persona 2: Primary users are CRM and lifecycle marketing leaders responsible for retention and engagement.
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?Most engagement platforms optimize campaigns. Pulse Intelligence optimizes customer behavior outcomes. Growth/Lifecycle PMs define the goal and constraints , review the customer intervention recommendation per user (as produced by the product) and hit execute or reject. Over time , once the product is mature and has credbility , PULSE executes the recommended action autonomously. Expected Product Inputs Pulse Intelligence relies on three categories of inputs. 1. First, behavioral event data from enterprise systems. This includes product views, searches, add-to-cart events, purchases, app opens, session activity, and engagement events such as clicks or message opens. 2. Second, communication history for each user across channels. This includes push notifications, emails, SMS, WhatsApp messages, in-app messages, frequency of sends, and user response patterns. 3. Third, business constraints defined by Lifecyle PMs. These include campaign goals such as revenue growth, retention improvement, or reactivation, along with constraints like maximum communication frequency, allowed channels, and rules around discount usage. Product Outputs The system produces decision-grade intelligence rather than dashboards or reports. The primary output is an intervention recommendation per user. This includes whether to send or suppress communication, which channel to use, and what type of intervention is most effective for each user. Each recommendation includes a predicted impact score, confidence level, expected incremental value, fatigue risk, and reasoning summary explaining why the decision was made. Secondary outputs include behavioral state scores such as purchase intent, churn risk, and engagement fatigue. The system also produces suppression recommendations to prevent harmful or low-value communication.
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?Lifecylce PMs managers currently cant answer the most important questions in their reviews: 1. The lifecycle PM running Braze who can't tell which step in her onboarding flow is over-sending 2. Growth PMs with actual P&L responsibility do not know if suppressing certain comms will lead to revenue loss , hence they continue to send low value comms , leading to low(er) ROI of their marketing budget 3. Multiple comms journeys targeting the same customer , yet the lifecyle PM does know the best intervention for each user . Even worse they dont know which customers are being blasted with comms when they should not be disturbed at all. (e.g. a Lifecycle PM ät a insuretech company . sending comms to purchase insurance , does not know that the user has a poor NPS wrt to a recently setlled claim)et. Key Pain point: "I'm rebuilding the same lifecycle in Braze every quarter, I can't tell which message is causing the opt-outs, and when leadership asks if our 9am send drove revenue I'm guessing. " "I dont know if a customer is being blasted by too much comms , I dont know if I am wasting my CAC on a customer who wasnt going to renew at all."
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.1. Intervention Recommendation: Pulse recommends amongst following actions for each user: a. send b. suppress c. wait d. best channel e. incentive vs no incentive 2. Suppression Intelligence: Predicts which users should NOT be contacted 3. Behavioral State Intelligence: Predicts: a. purchase intent b. churn risk c. fatigue d. engagement propensity Key AI opportunity include interpreting semantic event data , communication history , predicting recommendtion per user and explaining the reasoning behind the decision to earn trust from users.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.AI can help to understand user event stream , communication history , lifecyle behavior and generate intervention recommendation , suppression intelligence and customer behavioral state. AI can also help to generate best/personalised content for each communication , timing of sending the content , autonomous execution of the recommendation.
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.MVP will focus on establishing PULSE creditbility within marketers and prove that PULSE recommendations outweight traditional CRM communication. Which finally leads to increemental revenue , reduced marketing spends and discount wastage and less communication fatigue within customers. With this objective , below is the feature scope for MVP. 1. Intervention Recommendation: Pulse recommends amongst following actions for each user: a. send b. suppress c. wait d. best channel e. incentive vs no incentive 2. Suppression Intelligence: Predicts which users should NOT be contacted 3. Behavioral State Intelligence: Predicts: a. purchase intent b. churn risk c. fatigue d. engagement propensity Key AI opportunity include interpreting semantic event data , communication history , predicting recommendtion per user and explaining the reasoning behind the decision to earn trust from users. Not in MVP: copy generation, autonomous execution, budget optimization, multi-touch attribution.
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?1. Onboarding Flow 2. Integrations + Setup 3. Recommendation Action centre for each user where the user can auto-approve , observe or reject the actions for each user/user-group. 4. PULSE performance insights 5. Customer Behavioural insights Refer to the system architecture diagram in tab 2 for the full system workflow.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?Landing Page: https://preview--pulse-customer-intelligence.lovable.app/ Enterprise onboarding : https://preview--pulse-customer-intelligence.lovable.app/setup Action centre: https://preview--pulse-customer-intelligence.lovable.app/actions Product Impact Dashboard: https://preview--pulse-customer-intelligence.lovable.app/performance
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?Landing Page: https://preview--pulse-customer-intelligence.lovable.app/ Enterprise onboarding : https://preview--pulse-customer-intelligence.lovable.app/setup Action centre: https://preview--pulse-customer-intelligence.lovable.app/actions Product Impact Dashboard: https://preview--pulse-customer-intelligence.lovable.app/performance
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?1. EVENT_TAXONOMY_CLASSIFICATION_PROMPT = """ You are the semantic event taxonomy classifier for Pulse AI. Your job is to map raw enterprise event names into canonical behavioral meanings. This classification happens at taxonomy-build time or taxonomy-refresh time, not once per user. Allowed canonical meanings: - purchase_intent - exploration_intent - consideration_intent - churn_risk - conversion - engagement - inactivity - unknown Rules: - Infer semantic meaning from the event name and optional description only. - Do not infer user-level intent. - Do not predict outcomes. - Do not create business-specific operational labels. - Use "unknown" when the mapping is ambiguous. - Return strict JSON only. Expected JSON shape: { "mappings": [ { "raw_event_name": "add_to_cart", "canonical_meaning": "purchase_intent", "confidence": 0.94, "rationale": "Cart activity is a direct purchase-intent signal." } ] } """ EVENT_RUNTIME_REASONING_PROMPT = """ You are the Event Understanding reasoning layer for Pulse AI. Your role is to interpret a user's recent behavioral event stream after the system has already computed: 1. semantic event lookup signals, 2. embedding similarity scores, 3. deterministic rule candidates. You must resolve mixed signals and produce stable behavioral intent scores. Important cost and architecture constraints: - You are not the primary event classifier. - You are not called for every user by default. - You are used only for low-confidence, ambiguous, or high-value fallback cases. - Prefer the embedding scores unless the raw events clearly contradict them. - Do not generate marketing copy. - Do not recommend actions. - Do not decide whether to send a message. Return strict JSON only with this shape: { "purchase_intent": 0.0, "exploration_intent": 0.0, "churn_risk": 0.0, "behavioral_tags": ["tag"], "confidence": 0.0, "explanation": "Brief auditable rationale." } Scoring guidance: - All scores must be between 0 and 1. - Confidence should reflect signal agreement, recency, and ambiguity. - If purchase and fatigue/churn signals conflict, keep intent and risk separate. - If data is sparse, lower confidence rather than inventing certainty. 2. VOC_REASONING_SYSTEM_PROMPT = """ You are the Voice of Customer intelligence extraction layer for Pulse AI. Your role is to analyze customer communication across calls, chat logs, emails, support tickets, WhatsApp messages, and other customer messages. You are not a response-generation system. You are not a customer support agent. You are not deciding the next best action. Your only job is to extract universal behavioral communication signals from the customer's words and available acoustic metadata. Always infer these signals: - frustration_signal - urgency_signal - trust_signal - retention_risk - engagement_signal - escalation_risk Rules: - Scores must be floats between 0 and 1. - Use low confidence when communication is sparse, unclear, or mixed. - Keep trust_signal high only when the customer expresses confidence, patience, loyalty, or constructive engagement. - Treat cancellation, refund demands, repeated unresolved issues, or threats to leave as retention risk. - Treat repeated follow-ups, missed commitments, complaints, or angry tone as frustration and escalation signals. - Do not infer business-specific operational intents. - Do not recommend an action. - Do not generate a customer-facing response. - Do not reveal hidden chain-of-thought. - Return strict JSON only. Expected JSON shape: { "signals": { "frustration_signal": 0.0, "urgency_signal": 0.0, "trust_signal": 0.0, "retention_risk": 0.0, "engagement_signal": 0.0, "escalation_risk": 0.0 }, "behavioral_tags": ["tag"], "confidence": 0.0, "explanation": "Brief auditable rationale." } 3. NBA_EXPLANATION_SYSTEM_PROMPT = """ You explain Pulse AI next-best-action decisions for enterprise marketers and product users. The decision has already been computed by deterministic decisioning logic. Your job is to make the recommendation understandable, auditable, and trustworthy. You must explain: - why the selected action was chosen, - how the behavioral state influenced the recommendation, - how counterfactual actions compared, - how the business goal affected the ranking. Rules: - Do not change the selected action. - Do not recompute scores. - Do not invent missing metrics. - Use the provided counterfactual values directly. - Be concise and business-readable. - Do not reveal hidden chain-of-thought. - Avoid generic marketing copy. - Return plain text only.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?Product produces primary output as customer behavioural state , NBA and suppression predictions with a confidence score. Eval Metric - Event mapping accuracy - VoC sentiment/intent accuracy - Behavioral state accuracy - NBA action alignment - Hard rule violation rate - Suppression/service recovery recall
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?1. Accuracy of behavioural state prediction , NBA and suppression prediction for each user with different inputs (event stream x comms x CRM profile) is super critical. To test with atleast 20-50 users. Output accuarcy should be 90%+. 2.Strict JSON outputs , LLM outputs must match with the expected JSON output schema 3. System must follow defined hierarchy (for semantic event understanding) , refer to embedding rules , and only call LLM in case of mixed signals or ambiguity. 4. In case of mixed signals or ambiguity , LLMs must produce a low confidence score 5. Transcription quality and accuracy in case of different dialects or languages. Adversarial Testing 1. Test that customer protection rules always apply to the customer recommendation generated by the product. such as dont disturb the customer from 10 pm to 8 am , do not communicate with the customer when NPS is below threshold or active support ticket for more than 14 days is open or active escalation is ongoing. 2. Test that business policy and limits are always respected.
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?"event_taxonomy_classification": "gpt-5.1", "event_runtime_fallback": "gpt-5-mini", "voc_reasoning": "gpt-5.1", "nba_explanation": "gpt-5-mini", "audio_transcription": "gpt-4o-transcribe" 1. Event Semantic Meaning + Intent Use: gpt-5.2 or gpt-5.1 with reasoning.effort=medium But only for: new enterprise event taxonomy classification ambiguous event mappings periodic taxonomy refresh low-confidence edge cases Do not use it per user. Runtime should remain: raw event -> taxonomy lookup -> embeddings/rules -> LLM only if unresolved So your hierarchy is correct: embedding + rules first LLM reasoning only as fallback / taxonomy builder 2. VoC Agent Use: gpt-5.2 or gpt-5.1 with reasoning.effort=medium VoC has nuance: frustration, trust erosion, urgency, sarcasm, mixed sentiment, escalation risk. This is a better place to spend reasoning tokens. For audio: transcription: gpt-4o-transcribe or gpt-4o-mini-transcribe reasoning over transcript: gpt-5.1 / gpt-5.2 3. Behavioral State / NBA Explanation Layer Use: gpt-5-mini This is not deep reasoning. It is explanation generation from already-computed facts: state vector ranked actions counterfactuals constraints selected action So gpt-5-mini is the right default for cost efficiency.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.Persona Main Task Output JSON structure Expected Input Stucture Strict Rules on what NOT to do
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?1. Customer Protection Rules : Exclude customers when sending would do more harm than good. such as NPS below threshold , open support ticket for more than 14 days etc. These rules can be custom defined by the user and override all recommendations by the system. 2. Set brand and business limits that must be respected by the product at all times, such as : a. Quiet hours e.g. never disturb the customer from 10 pm to 8 am) b. Max allowed discount c. Max messages per user/day/week These rules should supersede any NBA prediction generated by PULSE.
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)Event mapping accuracy > 90% VoC semantic adequacy > 85% Hard constraint violation rate = 0% Recommendation Precision Rate > 80%
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?"Human is always in the loop"to finally execute or reject the action. This helps PULSE build trust. System produces every user recommendation with a confidence score, along with a simple to understand reasoning. a. User can auto approve all recommendations where confidence score is >80% b. observe where confidence score is between 60-80% c. and straightaway reject recommendations with < 60% confidence. b and c require qualitative assessment by the user , after reviewing the system generated explanation behind the recommendation.
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.1. EVENT_TAXONOMY_CLASSIFICATION_PROMPT = """ You are the semantic event taxonomy classifier for Pulse AI. Your job is to map raw enterprise event names into canonical behavioral meanings. This classification happens at taxonomy-build time or taxonomy-refresh time, not once per user. Allowed canonical meanings: - purchase_intent - exploration_intent - consideration_intent - churn_risk - conversion - engagement - inactivity - unknown Rules: - Infer semantic meaning from the event name and optional description only. - Do not infer user-level intent. - Do not predict outcomes. - Do not create business-specific operational labels. - Use "unknown" when the mapping is ambiguous. - Return strict JSON only. Expected JSON shape: { "mappings": [ { "raw_event_name": "add_to_cart", "canonical_meaning": "purchase_intent", "confidence": 0.94, "rationale": "Cart activity is a direct purchase-intent signal." } ] } """ EVENT_RUNTIME_REASONING_PROMPT = """ You are the Event Understanding reasoning layer for Pulse AI. Your role is to interpret a user's recent behavioral event stream after the system has already computed: 1. semantic event lookup signals, 2. embedding similarity scores, 3. deterministic rule candidates. You must resolve mixed signals and produce stable behavioral intent scores. Important cost and architecture constraints: - You are not the primary event classifier. - You are not called for every user by default. - You are used only for low-confidence, ambiguous, or high-value fallback cases. - Prefer the embedding scores unless the raw events clearly contradict them. - Do not generate marketing copy. - Do not recommend actions. - Do not decide whether to send a message. Return strict JSON only with this shape: { "purchase_intent": 0.0, "exploration_intent": 0.0, "churn_risk": 0.0, "behavioral_tags": ["tag"], "confidence": 0.0, "explanation": "Brief auditable rationale." } Scoring guidance: - All scores must be between 0 and 1. - Confidence should reflect signal agreement, recency, and ambiguity. - If purchase and fatigue/churn signals conflict, keep intent and risk separate. - If data is sparse, lower confidence rather than inventing certainty. 2. VOC_REASONING_SYSTEM_PROMPT = """ You are the Voice of Customer intelligence extraction layer for Pulse AI. Your role is to analyze customer communication across calls, chat logs, emails, support tickets, WhatsApp messages, and other customer messages. You are not a response-generation system. You are not a customer support agent. You are not deciding the next best action. Your only job is to extract universal behavioral communication signals from the customer's words and available acoustic metadata. Always infer these signals: - frustration_signal - urgency_signal - trust_signal - retention_risk - engagement_signal - escalation_risk Rules: - Scores must be floats between 0 and 1. - Use low confidence when communication is sparse, unclear, or mixed. - Keep trust_signal high only when the customer expresses confidence, patience, loyalty, or constructive engagement. - Treat cancellation, refund demands, repeated unresolved issues, or threats to leave as retention risk. - Treat repeated follow-ups, missed commitments, complaints, or angry tone as frustration and escalation signals. - Do not infer business-specific operational intents. - Do not recommend an action. - Do not generate a customer-facing response. - Do not reveal hidden chain-of-thought. - Return strict JSON only. Expected JSON shape: { "signals": { "frustration_signal": 0.0, "urgency_signal": 0.0, "trust_signal": 0.0, "retention_risk": 0.0, "engagement_signal": 0.0, "escalation_risk": 0.0 }, "behavioral_tags": ["tag"], "confidence": 0.0, "explanation": "Brief auditable rationale." } 3. NBA_EXPLANATION_SYSTEM_PROMPT = """ You explain Pulse AI next-best-action decisions for enterprise marketers and product users. The decision has already been computed by deterministic decisioning logic. Your job is to make the recommendation understandable, auditable, and trustworthy. You must explain: - why the selected action was chosen, - how the behavioral state influenced the recommendation, - how counterfactual actions compared, - how the business goal affected the ranking. Rules: - Do not change the selected action. - Do not recompute scores. - Do not invent missing metrics. - Use the provided counterfactual values directly. - Be concise and business-readable. - Do not reveal hidden chain-of-thought. - Avoid generic marketing copy. - Return plain text only.
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?Evolution and Changes 1. Added strict json format structures as output examples for every prompt , since some of the LLM outputs were viplating the recommended API contracts. 2. Added default instructions in case of ambiguity , to prevent incorrect outputs such as "Use "unknown" when the mapping is ambiguous""". Added additional instructions - Do not invent missing metrics. 3. Asked LLM to always publish confidence scores and always explain the REASONING for each recommendation. 4. Experimented with different models and reasoning effort for different modules calling LLMs , based on use-case and complexity , for optimum cost aware design. 5. Focussed on batch processing vs real time processing (when critical events such as add to cart are triggered) when making LLM calls , to manage cost (at a higher latency since the product user can wait a few seconds for the output). 6. Created Embedding vectors for within different modules for fast pattern matching first and called LLMs later when embedding vector do not retrun high confidence output. Clearly defined system hierarchy - raw event -> taxonomy lookup -> embeddings/rules -> LLM only if unresolved
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?Synthetic enterprise datasets generated across permutations of: • event streams • CRM profiles • communication history • business constraints • multilingual VoC • noisy instrumentation These datasets form the golden evaluation suite used for regression testing and prompt validation.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.Used AI to mimic user test input data (in the absence of a real enterprise client data) for a ecommerce company for diverse scenarios. Created a golden data set with 50 diverse user scenarios including ambiguos signals/comms history/CRM data , jarged voice transcriptions , model recommendations conflicting with user custom constraints. (forms the golden test data set) 1. Event taxonomy suite 2. VoC multilingual/noisy communication suite 3. Behavioral state suite 4. Mixed-user signal suite 5. NBA recommendation suite 6. Constraint conflict suite 7. Confidence band calibration suite 8. Explanation quality suite 9. Expected value sanity suite
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)1. Mixed or ambiguous user event signals 2. Scale : Enterprises have millions of customers and billions of user event streams.System doesnt loop/get stuck in case of large event data streams is critical 3. Jargled voice logs with different language or accents may generate poor transcripts leading to poor system outputs. So, this is aimportant adversarial test case. 4. Custom business policies/customer protection rules may directly conflict with AI outputs. Clear precedence to user-defined rules, in such scenarious is key. System must respect the golden rules
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?1. NBA Model Finetuning : The NBA layer over-selected send_product_recommendation / send_discount_offer even for low-Trust or high_fatigue users. It was not recommending suppression or service recovery enough for complaint/fatigue scenarios. Similarly , model was not differentiating between exploration intent and Add to cart intent. Tested various such scenarious and finetuned the model. 2.Added Event prompt should output conflict_summary, evidence_events, and data_quality_flags. to the event taxonomy prompt for imporved event classification. 3. Improved the VoC prompt to explicitly support multilingual/code-mixed inputs, transcript quality flags, semantic confidence, and evidence spans. The local dialect transcrition test failed , which got fixed woth prompt imporvement. 3. Asked to always explain its work - NBA explanation prompt should explain confidence band, delivery plan, expected value per comms, and constraint overrides.
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?Current Golden Suite Result Overall executable pass rate: = 98% PASS Some metrics can only be tested with real data and user behaviour such as CTR lift , revenue lift , conversion lift. Metric Result Status Recommendation action alignment 19 / 20 = 95% PASS Suppression / service recovery recall 9 / 10 = 90% PASS Hard constraint compliance 5 / 5 = 100% PASS Mixed-signal auto-approve guardrail 8 / 8 = 100% PASS Confidence field completeness 25 / 25 = 100% PASS Expected value sanity check 25 / 25 = 100% PASS
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?1. Multilingual voice of customer logs , led to poor trasncription accuracy > prompt improved to handle multilingual logs 2. Mixed user signals may happen - due to noisy event instrumenttation , out of order events , multiple users sharing one account. Added confidence nands and data quality flags in the prompt for event taxonomy. 3. Added mandatory explanation layer and added a eval "Explanation quality score" for LLM to generate meaningful reasoning and getting the LLM to judge its work.
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?"1. NBA Model Finetuning : The NBA layer over-selected send_product_recommendation / send_discount_offer even for low-Trust or high_fatigue users. It was not recommending suppression or service recovery enough for complaint/fatigue scenarios. Similarly , model was not differentiating between exploration intent and Add to cart intent. Tested various such scenarious and finetuned the model. 2.Added Event prompt should output conflict_summary, evidence_events, and data_quality_flags. to the event taxonomy prompt for imporved event classification. 3. Improved the VoC prompt to explicitly support multilingual/code-mixed inputs, transcript quality flags, semantic confidence, and evidence spans. The local dialect transcrition test failed , which got fixed woth prompt imporvement. 3. Asked to always explain its work - NBA explanation prompt should explain confidence band, delivery plan, expected value per comms, and constraint overrides."
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?Golden data set will auto run through a script (encoded in a python file) and outputs grading % will be shared for all new clients onboarded. Failures will be flagged to the human for review.
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?Golden data set to be updated for new failure scenarious/edge cases uncovered post launch or whenever there is a change of prompts or rules or data set taxonomy.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?Embedding vector , module wise API outputs (to match with expected output schema) and contracts amongst all modules are tested. Evals are added to grade outputs across each of the LLM models and orchestrator layers.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?We will need 3 critical artifacts 1. Entrprise training document for Lifecyle PMs on how to setup and use the product. 2. Another document on how the enterprise should connect their event , CRM and overall comms history to the product. 3. Lastly , what to expect from the product within 1 week , month and a quarter and how to define succes metrics.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?1. Ensure exhaustive setup to ensure system makes the right customer outcome predictions: Connect Enterprise data (event data) , databases , comms channels such as braze , moengage etc. Have a pre-sales executive/product ops POC to manage this piece. 2. Observe to solve the cold start problem: For each new enterprise client , to solve the cold start problem , the product will observe the enterprise data for last 7 days , run and store outputs in the embedding vector. This is only for product training and not execution. With synthetic/cold-start data, many recommendations will have low confidence engine because of limited historical evidence. Outcome confidence improves through real/simulated memory records. So this step is key for outcome accuracy. 3. PIlot : Define a control and test group across and run pulse on the test group and document the success metrics weekly. Define pilot scope. Recommended Pilot Scope Start narrow: 1 product area or lifecycle journey 5,000 to 50,000 users if data is available 2 to 4 channels 3 to 5 intervention types 1 primary business goal, such as conversion, retention, or revenue Prove pilot success and move to 100% rollout. Pulse initially operates in recommendation mode.Autonomous execution is enabled only after enterprise confidence is established.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?Will monitor pilot success and then move towards 100% rollout. The pilot , roughly spanning across 2 weeks , will start narrow with key success metrics defined. The goal of the first two weeks is not full autonomous campaign optimization. The goal is to prove that Pulse can identify better customer-level decisions than static journeys. By the end of two weeks, client should expect: - connected sample data from events, CRM, and communication history - mapped event taxonomy for major customer actions - customer behavioral state generation - Action Centre recommendations for a pilot cohort - suppression, service recovery, channel, timing, and content-type decisions - explanations for every recommendation - a first measurement readout against a control or historical baseline ## Recommended Pilot Scope - 1 product area or lifecycle journey - 5,000 to 50,000 users if data is available - 2 to 4 channels - 3 to 5 intervention types - 1 primary business goal, such as conversion, retention, or revenue
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?1 executive document is created for Lifeclye PMs and executives on what pulse does , how to use it , pilot scope . what to exoect from the pilot and success metrics from the product, . Additonally , another document is created for the data and engineering team outlining key integrations and setup tasks along with critical API endpoints such as CRM endpoints. Both these files are avilable in the git repo. Below are the links to these: 1. Executive Readout 2. Integration Setup for data and engineering team
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?Will monitor pilot success and then move towards 100% rollout. The pilot , roughly spanning across 2 weeks , will start narrow with key success metrics defined. The goal of the first two weeks is not full autonomous campaign optimization. The goal is to prove that Pulse can identify better customer-level decisions than static journeys. Refer to Pilot Scope defined above. Key outcome: • Behavioral State • NBA decision agreement against human judgement > 90% • Explanations • Baseline comparison • Incremental impact dashboard
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?1. Any enterprise data (events , CRM data) will only be shared through whitelisted APIs to pulse. Secret keys to be kept safe. Additionally, any data present in client's database/warehouse will be made available . To store user_id only , while omitting any PII user data. 2. Raw comms files will be purged after extraction , wil extract signals such as sentiment, urgency, retention risk, complaint flags, and evidence snippets, then expire raw text quickly. 3. Configurable data retention windows based on enterprise and use-case. 3. Workspace/tenant isolation by default. Separate schema per enterprise.TO always use tenant-specific vector namespaces, never the actual tenant name , very much like how user_id functions. Guiding principle is Pulse that should store behavioral intelligence, not unnecessary customer identity or tenant identity.
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?Yes. Privacy and storage concerns are handled. 1. PII data to be masked/purged. user_id to be referenced only. 2. CRM data to only be accessed using whitelisted API through secret KEY and proper encryption 3. Separate database for each client stored by only tenant_id 4. Custom Retention time as defined by client and use-case. Raw data to be purged after extracting key values such as sentiment , behavioral state etc.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?Business metrics that demonstrate value: - Decision Quality - Decision Win/Acceptance Rate - Incremental value per Decision - Revenue per 1,000 Decisioned Users - Reduced marketing spend wastage User Metrics to define MVP Success | Recommendation action alignment on golden suite | `> 85%` | | Hard rule violation rate | `0%` | | Suppression/service-recovery recall | `> 75%` | | Notification reduction | `10-20%` in eligible journeys | | CTR or conversion uplift | directional positive in pilot | ## Pilot Readout Questions - Which users did Pulse suppress that static journeys would have messaged? - Which high-intent users received a better channel or time recommendation? - Which users needed service recovery before marketing?
AI MetricsHow will you measure AI performance and accuracy?Eval Metric Target Event mapping accuracy >90% VoC sentiment/intent accuracy >85% Behavioral state accuracy >85% NBA action alignment >85% Hard rule violation rate 0% Suppression/service recovery recall >75% AI drift, data pipeline freshness Model latency/cost vector DB health API errors
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?Enterprise Support • In-product support • Dedicated Slack/Teams channel • Email support • P0 on-call escalation • Named Customer Success Manager • Published SLA by incident severity
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?Any downtime , API outages are categorized as P0 and trigger P0 alerts to both client and pulse. Golden dataset runs through a script whenever new data sources/model/prompt changes are done. Reduction in pass threshold generates amn alert along with failure test case. Bugs or feedback to be categorized as P0/P1/P2 Incidents with coresponding fix timelines with progress communication to the client.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?Evals to continually generate pass/fail metrics through a script , in addition to API downtime/outages etc. If any evaluation falls below threshold , trigger to be generated. Any bug fixes updates the product version and coresponding testcase to be added to the golden test suite.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?Evals to continually generate pass/fail metrics through a script , in addition to API downtime/outages etc. If any evaluation falls below threshold , trigger to be generated. Any bug fixes updates the product version and coresponding testcase to be added to the golden test suite.
Download the .xlsx ↓