← All capstone projects

AI Tools

FM Predict

Built by Dennis Beireis Cohort 9 Facilities management / predictive maintenance

FM Predict is an AI layer on top of an existing CAFM system that flags likely refrigeration failures before they become emergencies. It analyzes historical work orders, identifies repeat repair patterns and shortening intervals, and presents coordinators with prioritized alerts, evidence, and estimated savings so they can approve preventive actions quickly. The goal is to shift facilities operations from reactive repair tracking to foresight-driven maintenance across thousands of stores.

The problem

Retail facilities management is reactive by design: work orders are only created after something breaks. For an FM Coordinator managing 80 stores, refrigeration failures are always a surprise, always expensive, and always a store-manager escalation. Each missed failure means €500–2,000 in food spoilage plus an emergency vendor premium of 2–3x the planned rate — and the same unit often gets repaired three or four times before anyone notices the pattern. Refrigeration alone accounts for ~40% of total FM spend, yet there is no failure foresight, no repeat-offender flagging, and no predictive tooling.

The solution

FM Predict is an AI layer on top of the existing ServiceChannel CAFM system that flags likely refrigeration failures before they become emergencies. It analyzes historical work-order history per asset, detects repeat repair patterns and shortening intervals, and presents coordinators with prioritized alerts, evidence, and estimated savings. Each alert carries a risk score, the asset's repair history, a cost comparison (emergency vs. preventive), and a pre-filled preventive work-order draft the coordinator can approve with one click. It sits inside ServiceChannel with zero new tools, capped at five alerts per coordinator per week in a deliberate precision mode, shifting operations from reactive tracking to foresight-driven maintenance.

How it works

A nightly job analyzes ServiceChannel WO history across 500+ stores. When an asset matches a repeat-failure threshold — for example, ≥3 same-trade work orders in 18 months with decreasing intervals — the system calculates a failure-probability score and cost delta, then uses Claude (Haiku 4.5 in the deployed prototype) to generate a plain-English alert and pre-filled WO draft. Strict rules govern behavior: only same-trade repeats count, scheduled PM work orders are excluded, WO IDs are never invented, and thresholds are enforced (HIGH ≥75%, MEDIUM ≥50%, LOW ≥40%). Output is valid JSON, no RAG needed since Claude's long context handles 24 months of history in the API call. Live testing reproduced WO IDs and interval patterns exactly at 87% confidence, with a 97.6% objective pass rate (41/42 criteria) across six test cases; a human always approves before any WO is created.

Who it's for

The primary end users are FM Coordinators — each managing 50–150 stores and creating 20–50 work orders a week, working reactively today — for whom one missed failure is a costly, stressful escalation. Regional FM Managers own the budget and get a portfolio dashboard of active alerts, prevented incidents, and estimated savings. Within the retailer (Aldi, a 10,000+ store chain), the internal buyers are Regional FM Managers and the Head of FM, whose decision criteria are measurable cost reduction, risk mitigation, and minimal disruption to existing ServiceChannel workflows.

Why it matters

The global predictive maintenance market is projected to grow from ~$6B (2024) to ~$28B (2029), a ~36% CAGR, with refrigeration-specific predictive maintenance among the fastest-growing sub-segments due to food-safety and energy mandates. Retail FM is uniquely underserved: no sensors, no predictive tradition — yet years of structured ServiceChannel WO history form an unmatched data moat no external vendor can replicate. By converting emergency repairs into planned ones, FM Predict targets cost avoidance at scale. The launch is a phased 90-day pilot (120 stores, refrigeration only) with a KPI of ≥3 prevented failures and ≥€5,000 documented savings, expanding regionally then nationally, with human-in-the-loop approval preserving accountability throughout.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Dennis Beireis
Your Product:FM Predict – AI Predictive Maintenance Triage
Your Industry:Facilities Management / Retail
Date:12.05.2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Facilities Management / Retail. Aldi operates one of the largest discount supermarket chains globally with 10,000+ store locations requiring continuous infrastructure maintenance. FM is a critical cost driver — refrigeration alone accounts for ~40% of total FM spend.Please leave this area blank. This space is for the Instructor to provide you with feedback.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?Headwinds: rising energy costs, aging store infrastructure, technician shortages, reactive-only maintenance culture, limited FM visibility at store level. Tailwinds: rich ServiceChannel historical WO data, AI-driven predictive maintenance maturing rapidly, IoT sensor costs falling. Key FM software competitors: Corrigo, MaintainX, UpKeep, IBM Maximo. Why FM in retail is uniquely underserved by AI: unlike manufacturing (continuous sensor data) or IT (digital-native systems), retail FM is reactive by design — work orders are only created after a failure occurs. There is no tradition of predictive tooling, no sensor infrastructure, and no budget for standalone AI platforms. This makes the ServiceChannel WO history data an unusually high-value signal source: it is the only structured data that exists, and it is already rich enough to detect patterns — if something intelligent is layered on top.
What is the projected growth rate of your target market segment over the next 3-5 years?Global predictive maintenance market: ~$6B (2024) → ~$28B (2029), CAGR ~36%. Retail FM software segment growing ~12% YoY. Refrigeration-specific predictive maintenance is one of the fastest-growing sub-segments due to food safety regulations and energy mandates.
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Mature. Aldi operates 10,000+ stores globally with a well-established FM function. The FM operations are stable but reactive. AI/predictive capabilities represent the next major modernization wave — early movers will capture significant cost savings.
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)Revenue through grocery and general merchandise retail (B2C). FM is a pure cost center. Business case for FM Predict: cost avoidance — reducing emergency repair costs (2–3x planned maintenance), eliminating food spoilage from refrigeration downtime (~€500–2,000 per incident), and extending asset lifespan.
Who is your primary customer base (B2B, B2C, B2B2C)?B2C (end consumers). Internally, FM Predict serves B2B users: FM Coordinators (daily WO management), Regional FM Managers (portfolio oversight, budget), and indirectly Store Managers (uptime).
DifferentiatorsWhat are the key differentiators for your company?Aldi's differentiators: lean operations, cost discipline, standardized store formats, owned real estate. For FM Predict: the proprietary ServiceChannel WO history across 500+ stores is an unmatched data moat — years of asset-level repair patterns that no external vendor can replicate.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?Internal buyers: Regional FM Managers and Head of FM / Facilities Director. Decision criteria: measurable cost reduction, operational risk mitigation, and minimal disruption to existing ServiceChannel workflows.
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?Primary: FM Coordinators — manage 50–150 stores each, create 20–50 WOs/week in ServiceChannel, work reactively today. Secondary: Regional FM Managers — own the budget and KPIs. FM Coordinators are most impacted: one missed refrigeration failure = €500–2,000 food loss + emergency vendor premium + store manager escalation.
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?ServiceChannel is Aldi's CMMS for WO creation, vendor dispatch, invoice management, and asset tracking across 500+ stores. Current state: zero predictive layer — every repair is triggered by a failure event. FM Predict would sit as an AI layer on top of ServiceChannel data, not replacing it.
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)Primary persona: Sarah, FM Coordinator at Aldi Germany. Manages 80 stores across a regional cluster. Logs into ServiceChannel daily. Creates ~35 WOs/week, ~60% reactive. Biggest recurring pain: refrigeration failures at stores she didn't see coming — always a surprise, always expensive, always a store manager escalation.
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?CURRENT STATE (Reactive): 1. Store Manager calls at 7am: 'Refrigeration unit in Aisle 2 is warm, product at risk.' 2. Sarah logs into ServiceChannel, creates Priority 1 emergency WO. 3. Dispatches nearest approved refrigeration vendor. 4. Vendor arrives 6–18h later (emergency surcharge applies). 5. Repair completed — compressor replaced. 6. €1,200 food spoilage already written off. 7. Sarah notes: this same unit had 3 compressor WOs in the past 14 months. FUTURE STATE (FM Predict): 1. FM Predict flags Filiale 042, Asset REF-042-C2 two weeks before failure: '4th compressor repair pattern detected — 87% failure probability within 21 days.' 2. Sarah reviews the AI-generated preventive WO draft in ServiceChannel. 3. Approves with one click. 4. Vendor scheduled for Tuesday night (planned rate). 5. Compressor replaced before failure. 6. Zero food loss. Zero emergency premium. Zero store manager call.
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?1. [SEVERE / DAILY] No failure foresight: Sarah has zero visibility into which of her 80 stores' assets are trending toward failure. Every breakdown is a surprise. 2. [SEVERE / FREQUENT] Emergency cost premium: unplanned refrigeration repairs cost 2–3x the planned rate. Across 500+ stores this adds up to millions annually. 3. [MODERATE / FREQUENT] Repeat failures ignored: the same asset gets repaired 3–4 times before anyone notices the pattern. No system flags repeat offenders. 4. [MODERATE] Food loss liability: each refrigeration downtime event risks €500–2,000 in spoilage write-offs plus potential food safety exposure. 5. [LOW / OCCASIONAL] Manual WO creation under pressure: creating a Priority 1 WO while a store manager is panicking on the phone is stressful and error-prone.
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.1. [PRIORITY #1 — SELECTED] Repeat failure pattern recognition: AI analyzes ServiceChannel WO history per asset, detects refrigeration units with repeat repair patterns, generates a preventive WO draft with asset context, failure probability, and recommended action. Directly addresses pain points #1, #2, and #3. 2. [PRIORITY #2] NLP on WO description text: LLM extracts early warning signals from WO notes ('unit running warm', 'compressor cycling') to catch issues before they become emergencies. 3. [PRIORITY #3] AI-generated asset health summaries: before dispatching a vendor, FM Coordinator gets a one-paragraph AI summary of the asset's full repair history — reducing repeat diagnosis time.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.A) Failure pattern classifier on WO history (repair frequency, recency, asset age, trade category). B) NLP model on WO description text to extract early failure signals. C) Automated preventive WO generation with pre-filled vendor, priority, and asset context. D) Asset health score dashboard (RAG status: Green / Amber / Red) per store across 500+ locations. E) Proactive email/push alert to FM Coordinator when asset crosses risk threshold. F) Predictive cost model: shows estimated savings of preventive vs. emergency repair per asset. G) RAG-based asset manual lookup: vendor gets AI-generated repair guidance on arrival. H) Anomaly detection on invoice data: flags assets where repair costs are trending up abnormally.
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.RANKING by Impact × Feasibility: 1. [SELECTED ✓] Failure Pattern Classifier + Auto-WO Generator Impact: HIGH — directly prevents the most costly FM event (refrigeration emergency) Feasibility: HIGH — ServiceChannel WO history data already exists and is structured Output: Preventive WO draft auto-generated for FM Coordinator 1-click approval 2. Asset Health Score Dashboard (RAG status) Impact: HIGH — gives Regional FM Manager portfolio-wide visibility Feasibility: MEDIUM — requires UI build on top of classifier output 3. NLP on WO Description Text Impact: MEDIUM — catches earlier signals but lower precision than history patterns Feasibility: MEDIUM — requires NLP pipeline on unstructured text FOCUS: Solution 1. FM Predict analyzes ServiceChannel WO history per asset across 500+ stores, identifies refrigeration units matching repeat-failure patterns (e.g. ≥3 same-trade WOs within 18 months), calculates failure probability score, and auto-generates a preventive WO draft in ServiceChannel for FM Coordinator review.
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?1. DATA LAYER: Nightly analysis of ServiceChannel WO history per asset across 500+ stores. 2. DETECTION: Asset matches repeat-failure threshold (≥3 same-trade WOs in 18 months with decreasing intervals) → failure probability score + cost delta calculated. 3. ALERT GENERATION: Claude 3.5 Sonnet generates pre-filled preventive WO draft: asset ID, store, trade, vendor, scheduling window, cost comparison, reasoning. 4. FM COORDINATOR ACTION: Alert in ServiceChannel FM Predict tab → review evidence → 1-click 'Approve & Create WO'. 5. WO DISPATCH: Enters standard ServiceChannel dispatch workflow — vendor notified. 6. OUTCOME TRACKING: FM Predict logs whether action was taken and whether failure occurred → feeds model improvement. 7. REGIONAL VIEW: Regional FM Manager sees portfolio dashboard: active alerts, prevented incidents MTD, estimated savings, store risk ranking.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?4-screen navigation flow (zero new tools — all inside ServiceChannel): Screen 1 — Alert Inbox (new FM Predict tab in ServiceChannel nav): → Max 5 prioritized alerts/week per coordinator (Precision Mode) → Each card: risk % badge, asset ID, store, repair evidence summary, cost comparison (emergency vs preventive) → Actions: 'Approve WO' (1-click) | 'View details' Screen 2 — WO Approval Screen: → Pre-filled: store, asset, trade, suggested vendor (KälteTech GmbH ★4.8), suggested date/window (off-hours) → Cost comparison block: €3,200 (emergency, strikethrough) → €850 (preventive) → €2,350 saving → AI reasoning block: plain English explanation citing specific WO IDs → Actions: 'Approve & Create WO' | 'Edit WO details' | 'Dismiss alert' Screen 3 — Asset History View: → 4 metric cards: total WOs, total spend, asset age, risk score → Full repair timeline with colored dots (amber = planned, red = emergency, green dashed = pending preventive) Screen 4 — Regional Dashboard (FM Manager only): → 4 KPI cards: active alerts, prevented this month, estimated savings MTD, alert approval rate → Store risk ranking: top stores by risk score with RAG bar visualization → Read-only — no actions required Prototype link: https://fmpredict.lovable.app
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?FM Predict prototype demonstrates the full 4-screen AI user flow built in Lovable.dev: 1. Alert Inbox — AI-ranked alerts with risk scores, asset evidence, and cost comparison cards 2. WO Approval Screen — pre-filled WO draft with Claude's reasoning block and 1-click approval 3. Asset History View — full repair timeline for Filiale 042 / REF-042-C2 anchor example 4. Regional Dashboard — portfolio KPIs and store risk ranking for Regional FM Manager view AI interactions demonstrated: - Pattern detection output displayed as structured alert card - Claude reasoning explanation rendered in plain English with WO ID citations - Cost comparison (emergency vs preventive) calculated and displayed per alert - Success confirmation after WO approval ('WO PRV-2026-0042 created. Vendor KälteTech GmbH notified.') Prototype link: https://fmpredict.lovable.app
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?INITIAL MASTER PROMPT (Claude 3.5 Sonnet via Anthropic API): You are FM Predict, an AI assistant integrated into ServiceChannel for Aldi's Facilities Management team. Your role is to analyze Work Order (WO) history data for specific assets and generate predictive maintenance alerts when failure patterns are detected. Your outputs are used by FM Coordinators who manage 50–150 stores each. They are experienced FM professionals — not AI experts. They need clear evidence, not probability theory. OUTPUT FORMAT: Always respond in valid JSON: { 'alert_triggered': boolean, 'risk_score': integer (0-100), 'risk_level': 'HIGH' | 'MEDIUM' | 'LOW' | 'NONE', 'reasoning': string (max 3 sentences, plain English, cite specific WO IDs), 'pattern_type': 'repeat_failure' | 'age_based' | 'frequency_acceleration' | 'none', 'wo_draft': { asset_id, store_id, trade, priority, description, suggested_vendor, suggested_window, estimated_preventive_cost, estimated_emergency_cost, estimated_saving } | null } RULES: 1. Only flag SAME-TRADE repeat WOs. Different trades = unrelated failures. 2. Thresholds: HIGH ≥75%, MEDIUM ≥50%, LOW ≥40%. No alerts below 40%. 3. Max 5 alerts per coordinator per week — rank by risk score, return top 5. 4. Never invent WO IDs. Only cite IDs present in the input. 5. Scheduled PM WOs do not count as failure signals. 6. Data quality issues (gaps >6 months): flag and reduce confidence by 15%. 7. Tone: professional, factual. No 'urgent' or 'critical' — use risk scores. 8. Language: English only.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?Good output criteria — a high-quality FM Predict alert must meet ALL of the following: OBJECTIVE (measurable): 1. ACCURACY: Risk score supported by ≥2 verifiable data points from WO history input. 2. COMPLETENESS: All alert card fields present — risk score, asset ID, store, evidence, cost comparison, vendor, scheduling window. 3. PRECISION: Alert only triggered when confidence threshold met (≥75% HIGH, ≥50% MEDIUM, ≥40% LOW). 4. CLARITY: AI reasoning max 3 sentences, plain English, references specific WO IDs. 5. ACTIONABILITY: Pre-filled WO draft complete enough for 1-click approval — no missing required fields. 6. TONE: Professional, factual. No alarmist language. Confident but not overconfident. 7. FORMAT: Valid JSON matching defined schema — parseable without post-processing. SUBJECTIVE (human judgment, assessed by FM Coordinator persona): 1. TRUST: Reasoning is transparent enough to act on without opening ServiceChannel manually. 2. EFFORT REDUCTION: WO draft eliminates all manual lookup steps. 3. TONE FIT: Sounds like a professional FM tool, not a chatbot. 4. FALSE POSITIVE TOLERANCE: Alert feels conservative — approving it and finding nothing wrong would not destroy trust. 5. VENDOR SUGGESTION QUALITY: Suggested vendor matches historical preference for this asset type.
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?POSITIVE CASES (should trigger alert): + REF-042-C2: 3 compressor WOs in 14 months, decreasing intervals (7mo→4mo→6wk) → HIGH 87% + REF-117-A1: 2 condenser WOs in 9 months, last repair 11 weeks ago → MEDIUM 64% + Asset age >8 years + 2 same-trade WOs in past 12 months → LOW alert with age-based reasoning NEGATIVE CASES (should NOT trigger alert): - Single WO in past 18 months regardless of cost or severity - Multiple WOs but different trades (electrical + refrigeration = unrelated) - Multiple WOs but stable/increasing intervals (not accelerating) - Asset serviced preventively <4 weeks ago - WOs marked as scheduled PM in ServiceChannel EDGE CASES (human review flag): ± Brand new asset (<1 year) with 2 WOs → flag for manual review, do not auto-alert ± WOs from a vendor later removed from approved list → note data quality issue in reasoning ± WO history gap >6 months → flag data quality, reduce confidence by 15% ± 2 different failure types on same asset within 30 days → flag as potential systemic issue
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?Model: Claude 3.5 Sonnet (Anthropic) via API. Rationale: 1. STRUCTURED OUTPUT: Claude excels at generating consistent, structured JSON outputs (alert cards, WO drafts) with minimal hallucination — critical for a production FM tool where a bad alert triggers a real vendor dispatch costing €500+. 2. REASONING TRANSPARENCY: Claude's explanations are clear and auditable — essential for the 'show me the evidence' requirement (Markus: 'I want to see the pattern, not just a score'). 3. LONG CONTEXT: Claude handles long WO history logs (50+ repair records per asset) without performance degradation. 4. SAFETY: Anthropic's Constitutional AI reduces risk of confidently wrong outputs. Alternatives considered: - GPT-4o: Strong competitor but Claude's structured output consistency and reasoning clarity edge it out for this use case. - Rule-based system: Simpler but cannot handle natural language WO descriptions or adapt to new failure patterns without manual updates.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.Required inputs per API call: 1. asset_id (string) — unique asset identifier from ServiceChannel (e.g. 'REF-042-C2') 2. store_id (string) — Filiale number (e.g. '042') 3. store_location (string) — city/region (e.g. 'Dortmund Nord') 4. asset_type (string) — category (e.g. 'Refrigeration — Compressor Unit') 5. wo_history (array) — last 24 months of WOs for this asset, each containing: wo_id, date, trade, description, vendor, cost, resolution_time, priority_level 6. asset_age_years (float) — age of the asset 7. avg_emergency_cost (float) — historical average emergency repair cost for this asset type 8. avg_preventive_cost (float) — historical average preventive repair cost for this asset type 9. preferred_vendors (array) — approved vendors for this trade at this store 10. blackout_windows (array) — times when vendor access is restricted (store opening hours)
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?Optional inputs (improve output quality if available): 1. last_inspection_date (date) — date of last formal asset inspection; reduces false positives on recently inspected assets 2. asset_model (string) — manufacturer/model number; enables model-specific failure rate lookup 3. energy_consumption_trend (string: 'stable'|'rising'|'falling') — rising energy draw can be an early failure signal for refrigeration 4. store_temperature_zone (string) — ambient temperature context (e.g. heated vs unheated storage areas affect failure rates) 5. coordinator_override_history (array) — previous alerts the coordinator dismissed; enables learning from rejected alerts to reduce false positives
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)Objective criteria — a high-quality FM Predict output must meet ALL: 1. ACCURACY: Risk score supported by ≥2 verifiable data points from input WO history. 2. COMPLETENESS: All required alert fields present (risk score, asset ID, store, evidence, cost comparison, vendor, scheduling window). 3. PRECISION THRESHOLD: Alert only triggered at ≥40% confidence. HIGH ≥75%, MEDIUM ≥50%, LOW ≥40%. 4. CLARITY: Reasoning max 3 sentences, plain English, cites specific WO IDs from input. 5. ACTIONABILITY: WO draft fields complete for 1-click approval — no missing required ServiceChannel fields. 6. TONE: Professional, factual, no alarmist language. Uses risk scores, not adjectives like 'urgent' or 'critical'. 7. FORMAT: Valid, parseable JSON matching the defined output schema. No extra fields, no missing fields.
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?Subjective criteria — assessed by FM Coordinator persona during manual review: 1. TRUST: 'Does this alert make me more confident or does it feel like a black box?' Good output shows its reasoning clearly with WO evidence. 2. EFFORT REDUCTION: 'Could I approve this WO without opening ServiceChannel manually?' Good output eliminates lookup steps entirely. 3. TONE FIT: 'Does this sound like a tool my FM team would use, or like a chatbot?' Professional, not conversational. 4. FALSE POSITIVE TOLERANCE: 'If I approve this and nothing was wrong, would I lose trust in the system?' Output should feel conservative, not trigger-happy. 5. VENDOR SUGGESTION QUALITY: 'Is the suggested vendor the one I would have picked?' Good output matches historical vendor preference per asset type at this store.
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.MASTER PROMPT v1.2 (deployed in Lovable prototype — production) Verified: this prompt is live at https://fmpredict.lovable.app — see src/lib/predict.functions.ts in the Lovable codebase. Live test output (Test 6, May 2026) confirmed exact reproduction of WO IDs, interval pattern, and confidence score. PRD prompt = deployed prompt. ✅ MASTER PROMPT v1.2: You are FM Predict, an AI predictive maintenance system integrated into ServiceChannel for Aldi's Facilities Management team. Analyze the following Work Order history data for 3 assets and generate exactly 3 predictive maintenance alerts. INPUT DATA: Asset 1: REF-042-C2 | Filiale 042 Dortmund Nord | Refrigeration — Compressor Unit | Age: 6.5 years WO History: WO #4821 (Mar 2024, Compressor replaced, €1,200, P1), WO #5103 (Oct 2024, Compressor cycling abnormally replaced, €1,350, P2), WO #5687 (Mar 2025, Compressor + refrigerant leak emergency repair, €2,100, P1) Intervals: 7 months → 4 months → 6 weeks (accelerating) Emergency cost avg: €3,200 | Preventive cost avg: €850 Asset 2: REF-117-A1 | Filiale 117 | Refrigeration — Condenser Unit | Age: 4.2 years WO History: WO #5201 (Aug 2024, Condenser cleaning + repair, €980, P2), WO #5544 (May 2025, Condenser unit partial failure, €1,100, P2) Emergency cost avg: €2,800 | Preventive cost avg: €720 Asset 3: HVAC-203-B3 | Filiale 203 | HVAC — Air Handler Unit | Age: 7.1 years WO History: WO #4990 (May 2024, Air handler belt replaced, €520, P3), WO #5388 (Feb 2025, Air handler vibration issue repaired, €610, P3) Emergency cost avg: €1,900 | Preventive cost avg: €600 RULES: 1. Only flag SAME-TRADE repeat WOs. 2. Risk thresholds: HIGH ≥75%, MEDIUM ≥50%, LOW ≥40%. 3. Always cite the specific WO IDs from the input — never invent WO numbers. 4. Savings = emergency cost minus preventive cost.
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?Prompt iteration log: v1.0 → Initial release (Week 4) Changes from design-phase analysis: + Added: 'PM WOs do not count as failure signals' + Added: 'New assets <12mo with 2 WOs: flag for manual review' + Added: data_quality_flag field to output schema + Refined: reasoning instruction to 'max 3 sentences, cite specific WO IDs' + Added: 'JSON only — no prose before or after' v1.0 → v1.1 (Week 6 — after live prompt testing, 5 test cases) Issue identified: Test 5 (stable intervals, no preferred_vendor in input) returned suggested_vendor=null — leaves WO draft incomplete for 1-click approval. Fix: Added Rule 9 to Master Prompt v1.1: + Rule 9: 'If no preferred_vendor is available in the input, set suggested_vendor to "To be assigned by FM Coordinator" — never return null for this field when wo_draft is not null.' Result: WO draft always complete and actionable regardless of vendor data availability. v1.2 — Week 6 (live prototype deployment on fmpredict.lovable.app): CHANGE: Prompt embedded directly in Lovable prototype with fixed asset input data (REF-042-C2, REF-117-A1, HVAC-203-B3) to enable real Claude API validation. MODEL: Claude Haiku 4.5 via Anthropic API (production deployment). VALIDATION: Live test confirmed — reasoning output cited WO #4821, #5103, #5687 correctly, interval pattern '7 months → 4 months → 6 weeks' reproduced accurately, confidence 87%, tone professional. ✅ NEXT: v1.3 planned post-pilot — add vendor rating filter (prefer ≥4.5★ vendors where multiple options available).
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?Data sources: 1. PRIMARY: ServiceChannel WO history API — structured JSON export of all WOs per asset for past 24 months. Fields: wo_id, date, trade, description, vendor_id, cost, status, priority_level, resolution_time_hours. 2. SECONDARY: ServiceChannel asset registry — asset_id, asset_type, installation_date, store_id, store_location. 3. TERTIARY: ServiceChannel vendor registry — vendor_id, name, rating, approved_trades, approved_stores. Data preparation: - Filter: exclude WOs with status 'CANCELLED' or 'DUPLICATE' - Normalize: convert all costs to EUR, all dates to ISO 8601 - Flag: mark WOs with priority_level 'PM' (planned maintenance) — excluded from failure signal detection - Quality check: flag assets with WO history gaps >6 months - No RAG implementation in MVP — all context passed directly in API call (Claude long context handles 24-month WO history within token limits)
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.Typical example inputs and expected outputs: EXAMPLE 1 — True Positive (High Risk): Input: REF-042-C2, 3 compressor WOs in 14 months (Mar 2024, Oct 2024, Mar 2025), intervals 7mo→4mo, last repair 6 weeks ago, asset age 6.5 years Expected: alert_triggered=true, risk_score=87, risk_level=HIGH Good output: reasoning cites WO #4821/#5103/#5687, wo_draft complete with KälteTech GmbH as vendor, Tuesday off-hours window, €850 preventive vs €3,200 emergency EXAMPLE 2 — True Negative: Input: REF-089-A1, 1 compressor WO in 18 months, no repeat pattern, asset age 3 years Expected: alert_triggered=false, risk_score=12, risk_level=NONE, wo_draft=null Good output: silent — no false positive generated EXAMPLE 3 — Medium Risk: Input: REF-117-A1, 2 condenser WOs in 9 months, last repair 11 weeks ago Expected: alert_triggered=true, risk_score=64, risk_level=MEDIUM Good output: amber alert, reasoning notes 2-WO pattern with accelerating frequency EXAMPLE 4 — Age-Based Risk: Input: HVAC-203-B3, 1 WO but asset age 9.5 years (near end of lifespan) Expected: risk_score=41, risk_level=LOW, pattern_type=age_based Good output: reasoning explicitly distinguishes age-based vs repeat-pattern signal
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)Edge cases and negative cases for evaluation: EDGE CASE 1: Mixed trade WOs Input: Asset with 2 electrical WOs + 1 refrigeration WO in 12 months Expected: No alert — different trades are unrelated failures Risk: Model incorrectly combines trades → false positive EDGE CASE 2: Brand new asset with early failures Input: Asset installed 8 months ago, 2 WOs already Expected: data_quality_flag set, no auto-alert, flag for manual review Risk: Model triggers alert on infant mortality pattern which has different root cause EDGE CASE 3: Data gap in history Input: Asset with 2 WOs but 8-month gap between them (data migration artifact) Expected: data_quality_flag='History gap >6 months detected', confidence reduced 15% Risk: Model treats gap as clean history and overestimates reliability EDGE CASE 4: Stable intervals (not accelerating) Input: Asset with 3 WOs but evenly spaced 8 months apart Expected: MEDIUM or LOW alert only — stable pattern, not accelerating Risk: Model triggers HIGH alert purely on WO count without considering interval trend NEGATIVE CASE: PM-only WOs Input: Asset with 3 WOs all marked priority_level='PM' Expected: No alert — all scheduled maintenance, no failure signals Risk: Model counts PM WOs as failure signals → false positive
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?Manual review results — Week 6 testing (May 2026) PHASE 1 — Simulated testing (5 cases, Master Prompt v1.0/v1.1): TEST 1 — True Positive (REF-042-C2, HIGH): risk_score=87, risk_level=HIGH ✅ | 7/7 objective criteria PASS TEST 2 — True Negative (REF-089-A1): alert_triggered=false ✅ | 7/7 PASS TEST 3 — Mixed Trade (REF-155-C1): alert_triggered=false, trades correctly separated ✅ | 7/7 PASS TEST 4 — New Asset Edge Case (REF-301-B1): manual review flag set, warranty review recommended ✅ | 7/7 PASS TEST 5 — Stable Intervals (REF-203-D2): risk_level=MEDIUM (not HIGH) ✅ | 6/7 PASS — suggested_vendor=null issue identified → resolved in v1.1 Simulated phase result: 34/35 objective criteria passed (97%). v1.1 resolves identified gap. PHASE 2 — Live prototype test (fmpredict.lovable.app, Master Prompt v1.2, Claude Haiku 4.5): TEST 6 — Live Reasoning Validation (REF-042-C2, Screen 2): Expected: WO #4821/#5103/#5687 cited, interval pattern reproduced, 87% confidence, professional tone Actual output: 'Asset REF-042-C2 has had 3 compressor repairs in the past 14 months (WO #4821, #5103, #5687). The average interval between failures is decreasing: 7 months → 4 months → last repair 6 weeks ago. Pattern matches 89% of assets that experienced a 4th failure within 30 days. Confidence: 87%' Result: ✅ ALL criteria met — WO IDs correct, interval pattern exact, confidence score consistent, tone professional, no alarmist language OVERALL: 6/6 tests passed. 41/42 objective criteria passed (97.6%). All identified issues resolved in subsequent prompt versions.
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?Automated evaluation plan — updated after Week 6 live testing: MVP (manual, pilot phase): - FM Coordinator scores each alert outcome after 30 days - Metrics tracked: precision, recall (sampled), approval rate, false positive rate - Baseline from testing: 97.6% objective pass rate (41/42 criteria, 6 test cases) Post-MVP (automated LLM grader): - Separate Claude call evaluates each output against 7 objective criteria - JSON schema validation scripted (100% automated) - Minimum 20-input test set (expanded from 6 live/simulated tests) - Nightly run; flags <80% score for human review - Re-run after every prompt version change Known gap from testing: vendor field completeness requires preferred_vendors in input — addressed in v1.1/v1.2. Monitor in production.
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?Edge cases identified during design phase (to be validated in testing): 1. Mixed-trade WOs on same asset — risk of incorrectly combining unrelated failures 2. Brand new assets (<12 months) with early repeat failures — infant mortality pattern differs from wear-out pattern 3. WO history gaps >6 months — data migration artifacts could create false clean periods 4. Stable vs accelerating failure intervals — WO count alone insufficient; interval trend must be analyzed 5. PM WOs miscounted as failure signals — requires explicit PM flag filtering 6. Vendor removed from approved list mid-history — historical WOs from that vendor remain in data 7. Store temporary closure (renovation) — WO gap during closure not a data quality issue Mitigation for each: handled in Master Prompt rules and data preparation filters. Edge cases 1–5 explicitly addressed in prompt rules.
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?Prompt version change log: v1.0 — Week 4 (design-phase analysis): + Added: 'PM WOs do not count as failure signals' + Added: 'New assets <12mo with 2 WOs: flag for manual review' + Added: data_quality_flag field to output schema + Refined: reasoning instruction — 'max 3 sentences, cite specific WO IDs' + Added: 'JSON only — no prose before or after' v1.1 — Week 6 (simulated testing, Test 5 finding): ISSUE: suggested_vendor=null when no preferred_vendor in input. FIX: Rule 9 — 'never return null for suggested_vendor when wo_draft is not null'. VALIDATION: Re-tested Test 5 ✅ v1.2 — Week 6 (live prototype deployment, fmpredict.lovable.app): CHANGE: Prompt embedded in Lovable with fixed asset input data for real Claude API validation. MODEL: Claude Haiku 4.5 (production). VALIDATION: Live Test 6 — reasoning cited WO #4821/#5103/#5687 correctly, interval pattern '7mo→4mo→6wk' exact, confidence 87%, tone professional ✅ NEXT: v1.3 post-pilot — add vendor rating filter (prefer ≥4.5★ vendors where multiple available).
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?Evaluation method: hybrid (human + model grader). MVP phase (pilot, 90 days): - Human review: FM Coordinator scores each approved alert outcome after 30 days - Criteria tracked: alert precision (prevented failures / total alerts), alert recall (sampled — failures on non-alerted assets), approval rate, false positive rate - Review cadence: weekly during pilot, monthly post-pilot Post-MVP (automated): - LLM-as-judge: separate Claude call evaluates each output against objective criteria - Script: JSON schema validation automated (100% coverage) - Diverse test set: minimum 20 inputs covering all positive, negative, and edge case categories - Scale: automated grader runs nightly on production outputs; flags outputs scoring <80% for human review
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?Evaluation frequency: During pilot (Weeks 1–12 post-launch): - Weekly: manual review of all alerts issued that week (small volume, high scrutiny) - After each prompt version change: full test suite re-run before deploying updated prompt Post-pilot (steady state): - Monthly: automated grader re-run on production outputs from past 30 days - Quarterly: full manual review with FM Coordinator feedback session - Trigger-based: immediate re-evaluation if false positive rate exceeds 20% in any week or if a high-risk alert (≥75%) is overridden by coordinator 3+ times in one month
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?Technical readiness checklist: ✓ Claude Haiku 4.5 API configured and tested via Lovable prototype ✓ ServiceChannel WO history API export pipeline validated ✓ Data preprocessing pipeline tested on sample data ✓ JSON schema validation (Zod) in place ✓ Alert delivery into ServiceChannel FM Predict tab confirmed ✓ 5-alert/coordinator/week cap enforced at application layer ✓ Rollback plan: if false positive rate >25% week 1, disable alert generation ✓ Monitoring dashboard: alert volume, approval rate, error rates tracked daily ✓ Master Prompt v1.2 deployed and validated — 97.6% pass rate on 6 test cases ✓ Live prototype: https://fmpredict.lovable.app (real Claude API confirmed)Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?Organizational readiness for 90-day pilot (120 stores, NRW region): ✓ Regional FM Manager (Markus Weber type) briefed — pilot KPI agreed: ≥3 prevented refrigeration failures in 90 days ✓ 3 FM Coordinators trained on FM Predict Alert Inbox and 1-click WO approval flow (30-min walkthrough) ✓ Store Managers in pilot region notified of potential increase in preventive vendor visits ✓ Vendor KälteTech GmbH briefed on potential increase in preventive dispatch volume ✓ IT/ServiceChannel admin: FM Predict tab enabled for pilot coordinator accounts only ✓ Legal/compliance review completed (see Row 52) ✓ Escalation path defined: FM Predict alerts do not bypass coordinator approval — human always in the loop Documentation complete: - FM Predict user guide (1-page, plain English) - FAQ for FM Coordinators: 'What is FM Predict and what do I do with alerts?'
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?Launch approach: phased pilot → regional rollout → national rollout. PHASE 1 — Pilot (Months 1–3): - Scope: 120 stores, NRW region, refrigeration assets only - Users: 3 FM Coordinators + 1 Regional FM Manager - Access: FM Predict tab enabled for pilot accounts only in ServiceChannel - Alert volume: max 5 alerts/coordinator/week (precision mode) - KPI: ≥3 prevented refrigeration failures, ≥€5,000 in documented savings PHASE 2 — Regional Rollout (Months 4–6, pending pilot success): - Scope: all Aldi Germany regions, refrigeration + HVAC - Users: all FM Coordinators and Regional FM Managers - Alert volume: scaled based on pilot precision calibration PHASE 3 — National Rollout (Month 7+): - Scope: all 500+ stores, all asset types - Features: Regional Dashboard enabled, mobile alerts, automated ROI reporting
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?Scale readiness plan: Pilot → regional scale (120 → 2,000+ stores): - API cost model validated: Claude 3.5 Sonnet API calls per asset per night estimated at ~$0.02/asset. At 10,000 assets across 500 stores: ~$200/night API cost — well within FM cost savings ROI. - ServiceChannel API rate limits: WO history export batched in nightly jobs to stay within rate limits - Alert volume management: coordinator alert cap (5/week) enforced regardless of scale — precision maintained - Monitoring: alert volume, API latency, error rates tracked via application dashboard - Infrastructure: stateless API design — scales horizontally without architectural changes - Rollback: prompt version pinning ensures consistent output regardless of Anthropic model updates
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?Internal launch communication assets (no external marketing — internal tool): 1. FM Predict One-Pager: 'What FM Predict does, why it matters, what to do when you get an alert' — distributed to all pilot FM Coordinators before go-live 2. 2-min demo video (screen recording of https://fmpredict.lovable.app): shared with Regional FM Managers before pilot 3. FAQ document: 10 most common questions from FM Coordinators during training ('What if I approve an alert and the asset is fine?' / 'Can I dismiss an alert?' / 'How many alerts will I get?') 4. Pilot results template: standardized monthly report format for Regional FM Manager to report prevented incidents and savings to Head of FM 5. ServiceChannel release note: brief in-app notification when FM Predict tab is enabled for pilot accounts
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?Internal stakeholder communication plan: WEEKLY (during pilot): - FM Coordinators receive weekly alert summary email: alerts issued, approved, dismissed, outcomes logged - Regional FM Manager receives weekly pilot metrics: prevented incidents, estimated savings, approval rate MONTHLY: - Head of FM / Facilities Director receives 1-page pilot summary: KPI status vs target, cost savings documented, recommendation for Phase 2 - IT/ServiceChannel admin: API usage, error rates, any technical issues flagged PILOT CONCLUSION (Day 90): - Full pilot debrief with Regional FM Manager and FM Coordinators - Go/no-go recommendation for Phase 2 regional rollout presented to Head of FM - Pilot ROI report: total savings documented, projected annual savings at full scale
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?Data privacy and protection protocols: 1. DATA SCOPE: FM Predict processes only operational FM data (asset IDs, WO history, costs, vendor names). No personal employee data, no customer data, no financial PII. 2. DATA RESIDENCY: ServiceChannel WO data remains in existing ServiceChannel infrastructure. Only anonymized asset/WO records are sent to Anthropic API — no store employee names, no personal identifiers. 3. ANTHROPIC API: Data processed under Anthropic's enterprise data processing agreement. API calls are not used for model training (zero data retention policy confirmed for enterprise tier). 4. GDPR COMPLIANCE: WO data contains no personal data as defined by GDPR — asset IDs and operational records are not personally identifiable. Legal review completed and confirmed compliant. 5. DATA RETENTION: Alert outputs retained in ServiceChannel for 24 months (consistent with existing WO retention policy). Raw API call logs retained 90 days then deleted. 6. ACCESS CONTROL: FM Predict alerts visible only to assigned FM Coordinators for their store cluster and their Regional FM Manager. No cross-regional data visibility.
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?Policy, compliance, and risk protocols: 1. HUMAN IN THE LOOP: FM Predict alerts are recommendations only — no WO is created without explicit FM Coordinator approval. System cannot autonomously dispatch vendors. 2. AUDIT TRAIL: Every alert generated, approved, edited, or dismissed is logged with timestamp and coordinator ID in ServiceChannel. Full audit trail for compliance review. 3. VENDOR CONTRACTS: FM Predict does not modify or create new vendor contracts. All dispatched WOs use existing approved vendor agreements in ServiceChannel. 4. AI DISCLOSURE: FM Coordinators and vendors are informed that WO recommendations are AI-generated. Alert cards display 'AI-generated recommendation' label. 5. CONTENT MODERATION: Not applicable — FM Predict generates structured operational data (JSON), not open-ended text content. 6. FOOD SAFETY REGULATIONS: FM Predict supports (not replaces) Aldi's food safety compliance requirements. Refrigeration failure prevention directly supports HACCP compliance obligations. 7. LIABILITY: FM Coordinator approval is the point of accountability. AI recommendation does not transfer liability — coordinator approving the WO takes operational responsibility.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?User and business success metrics: USER METRICS (FM Coordinator): - Alert approval rate: target ≥70% (indicates alerts are relevant and trustworthy) - Time to approve: target <2 minutes per alert (validates 1-click flow efficiency) - False positive rate: target <20% (alerts approved where no failure occurred within 30 days) - Coordinator satisfaction: quarterly NPS survey, target ≥7/10 BUSINESS METRICS (Regional FM Manager / Head of FM): - Prevented refrigeration failures: target ≥3 in 90-day pilot (user-defined KPI from Markus interview) - Cost savings (emergency vs preventive delta): target ≥€5,000 in pilot period - Emergency WO rate: target 10% reduction vs baseline in pilot region after 90 days - Food spoilage incidents: target 15% reduction in refrigeration-related spoilage write-offs - Pilot ROI: target >5x (savings vs FM Predict development + API cost)
AI MetricsHow will you measure AI performance and accuracy?AI performance metrics: 1. PRECISION: ≥80% HIGH alerts where asset failed within 30 days. 2. RECALL: ≥60% of actual failures preceded by FM Predict alert (sampled). 3. JSON VALIDITY: 100%. 4. THRESHOLD ADHERENCE: 100%. 5. WO ID ACCURACY: 100% — no hallucinated IDs. 6. REASONING COMPLETENESS: ≥95% of HIGH alerts cite ≥2 WO IDs. 7. API LATENCY: <3 seconds per asset. Baseline from testing (v1.2, 6 test cases): 97.6% objective pass rate (41/42 criteria). Live prototype validation: WO ID citation 100%, interval pattern reproduction 100%, tone compliance 100%. All identified gaps resolved in v1.1/v1.2.
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?Support channels for FM Predict pilot: 1. PRIMARY: FM Coordinator questions handled by the FM Predict product owner (Dennis Beireis) during pilot — direct Slack channel or email. 2. ESCALATION: Technical issues (alert not appearing, WO draft missing fields) escalated to IT/ServiceChannel admin within 4 hours during business hours. 3. FALSE POSITIVE REPORTING: 'Dismiss alert' button in FM Predict UI logs dismissal reason (dropdown: 'Asset was just serviced' / 'Pattern seems incorrect' / 'Other'). Product owner reviews dismissed alerts weekly. 4. OWNERSHIP: FM Predict product owner (Dennis Beireis) is accountable for alert quality, API uptime, and pilot KPI delivery. 5. POST-PILOT: Support transitions to standard FM IT helpdesk with FM Predict runbook provided.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?Feedback workflow: 1. IN-PRODUCT: 'Dismiss alert' dropdown captures dismissal reason. 'Approve WO' logs approval. Both feed directly into weekly metrics dashboard. 2. WEEKLY REVIEW: Product owner reviews all dismissals and any coordinator feedback messages. Patterns (e.g. 3+ dismissals citing 'asset just serviced') trigger prompt adjustment. 3. MONTHLY RETROSPECTIVE: 30-min call with FM Coordinators and Regional FM Manager. Structured feedback: What worked? What was frustrating? What would you change? 4. BUG REPORTING: Critical issues (alert created wrong vendor, cost comparison incorrect) reported via Slack to product owner → investigated within 24 hours → hotfix deployed within 48 hours. 5. PRIORITIZATION: Feedback triaged as: Critical (fix immediately) / High (next prompt iteration) / Low (post-pilot backlog). All feedback logged in a shared Notion doc accessible to all stakeholders.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?Monitoring and logging plan: APPLICATION MONITORING: - Alert volume per coordinator per week (enforces 5-alert cap) - Approval rate and dismissal rate — tracked daily during pilot - API error rate and latency — alerted if >2% error rate or >5s average latency - JSON schema validation failures — alerted on any invalid output AI QUALITY MONITORING: - Weekly: manual review of all generated reasoning blocks for hallucinated WO IDs - Weekly: spot-check 3 random alerts against source WO data for accuracy - Monthly: automated LLM grader evaluates sample of outputs against objective criteria OUTCOME TRACKING: - Each alert logged with: asset_id, risk_score, coordinator_action (approved/dismissed/edited), outcome_after_30_days (failure_occurred: yes/no) - Outcome data feeds precision and recall calculation monthly DASHBOARD: Real-time metrics visible to product owner and Regional FM Manager in Regional Dashboard screen.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?Continuous improvement process: PROMPT IMPROVEMENT: - Every 4 weeks: review false positives and dismissed alerts → identify prompt gaps → release new prompt version - Version history maintained in PRD (Row 36) with full change log - A/B testing post-pilot: test prompt v1.x vs v2.0 on 20% of assets before full rollout MODEL UPDATES: - Monitor Anthropic model release notes for Claude updates - Re-run full evaluation test suite on any new model version before switching - Pin to specific model version (claude-sonnet-4-20250514) to prevent unexpected output changes FEATURE EXPANSION: - Quarterly product review: new asset types (HVAC, electrical) added based on pilot data and FM team feedback - Post-pilot: evaluate NLP on WO description text as additional signal source - Year 1 target: expand from refrigeration-only to full FM asset portfolio across all 500+ stores LEARNING LOOP: - Alert outcome data (failure_occurred: yes/no) fed back into pattern threshold calibration every quarter - Goal: continuously improve precision from ≥80% (pilot target) toward ≥90% (steady state target)
Download the .xlsx ↓