E-commerce
Beneath
Beneath is a quality investigation agent for weekly business reviews in a retail distribution center. It starts from a deterministic dashboard layer to build trust in a legacy environment, then adds an agentic investigator and composer that generate speaker notes, callouts, and a ready-made presentation deck. The workflow includes review, feedback, regeneration, and a scheduled weekly run for operational use.
The problem
In a legacy retail distribution center, the ICQA Manager authors the Weekly Business Review that leadership relies on for site quality. The current rules-based platform deterministically picks the worst-gap process path as the deep-dive focus — and in a meaningful share of weeks, that pick misses the real story. When it does, the manager must do analytical work the tool cannot: spotting a long-running out-of-tolerance pattern, an outsized single-shift anomaly, or a department drifting quietly below the threshold, then rewriting the narrative to match. That edit-required cycle adds up to 2 hours and is the single largest variable cost in the workflow — materially worse as the site consolidates from two ICQA Managers to a solo-coverage model.
The solution
Beneath is a quality investigation agent for weekly business reviews. It starts from a deterministic dashboard layer to build trust in an environment that is not yet AI-native, then adds an agentic investigator and composer that generate secondary callouts, speaker's notes, and a ready-made presentation deck. The deterministic worst-gap logic remains the spine of the deep-dive; the agent's editorial authority is deliberately bounded to secondary callouts and speaker's notes that enrich the narrative with cross-week judgment. The workflow includes review, per-callout veracity flags, feedback, regeneration, and a scheduled weekly run — collapsing the 1–2 hour edit cycle into a roughly 5-minute review cycle in most weeks.
How it works
Beneath runs a two-phase Claude Opus 4.7 conversation on a weekly cron (Monday 9 AM Central). An investigator phase reads the deterministic deck's payload, then loops through investigation-tier API endpoints — every reason code, workstation, team member, and freight category — building structured candidate observations with confidence scores and citations. A composer phase writes the final callouts and speaker's notes in the voice of an experienced site quality manager. A layered guardrail stack gates every claim: a forbidden-phrase guard applied at three points to keep output observational (never prescriptive), a bare-code regex that rejects unlabeled reason codes, and a confidence floor at 30 that drops weak observations. An advisory veracity layer — deterministic grounding plus a best-effort LLM judge — surfaces per-callout flags in the ReviewPanel, and a run-level eval gate requires all factual-accuracy checks plus 80% of semantic checks to pass; a failed gate triggers graceful degradation to a deterministic-only deck.
Who it's for
Beneath is a B2E internal product inside a Fortune 50 omnichannel mass retailer. The primary user is the ICQA Manager at a single distribution center, who owns site-level quality reporting and manages a team of 30–40. Secondary users are Operations Managers and Senior Operations Managers, who gain self-serve access to current reports, and tertiary consumers are the Site Director and Operations Director. The economic buyers are the DC operations leadership chain — Site Director and Operations Director — who sponsor tooling that reduces reporting burden and improves review consistency. Regional VP and above are explicitly out of scope.
Why it matters
The site is piloting a future network staffing model, consolidating two ICQA Managers into one; going from an edit-required reporting cycle to a 5-minute review cycle is the difference between a workable solo role and one absorbed by reconciliation and authoring. The business value is realized indirectly: manager hours redirected to floor work and coaching, faster leadership decision cycles, and quality issues caught earlier. The North American warehouse operations software market is projected to grow at roughly 10–14% CAGR through 2028, with AI-assisted reporting growing faster. Beneath is now built and published to production, deliberately observational rather than prescriptive — the appropriate posture for a supply chain building trust with AI before any prescriptive surface earns its place in a future phase.
The workflow
The PRD
| PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0 | ||||||
|---|---|---|---|---|---|---|
| Your Name: | Niya Shorter | |||||
| Your Product: | BIneath | |||||
| Your Industry: | Retail Supply Chain | |||||
| Date: | June 15, 2026 | |||||
| 4D Method | AI PRD | Instructor Feedback | ||||
| Phase | Activity | Theme | Topic | Key Question(s) | Your Response Include external links to visuals/prototypes as required. | |
| DISCOVERY | Understand your market, business, product & user context | Business Value Map | Market Attractiveness | What industry is your business in? (ie Financial services, Healthcare, Education, etc)? | Retail Supply Chain & Logistics, specifically the inbound and outbound operations of large-format omnichannel mass retailers that run Distribution Centers (DCs) to flow inventory from vendors to stores and to fulfill direct-to-guest digital orders. BIneath sits inside the DC Operations slice of this industry, supporting Inventory Control & Quality Assurance (ICQA) Managers as the primary persona, with Operations Managers (OMs), Senior Operations Managers (SOMs), Site Directors, and Operations Directors as secondary and tertiary stakeholders. | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors? | Headwinds: persistent labor scarcity and turnover in DC frontline roles, compressed margins from promotional pressure and tariffs on imported goods, and a tightening expectation from corporate leadership that site-level operations report on quality metrics with the same rigor as financials. Defect Rate (cancellations divided by demand volume) is the dominant lens for site quality, and leadership cadences (Daily Business Reviews and Weekly Business Reviews) have hardened around it. A specific headwind for legacy DCs: many sites are 25–35 years old physically and technologically, with operational data fragmented across legacy systems (homegrown card-based tools, Splunk dashboards, 3D WMS, SharePoint workbooks filled out by hand) and modern enterprise data platforms in parallel. The reconciliation burden between those two worlds falls on the ICQA Manager. Tailwinds: active WMS modernization underway across the industry, maturing enterprise data infrastructure, and growing acceptance of LLM-assisted analysis in enterprise environments under zero-data-retention contracts. The combination means the substrate for agentic operations products now exists where it did not three years ago, and the modernization moment is an opening to introduce AI-native tooling alongside the new WMS rather than as a retrofit later. Competitors for the leadership-reporting use case BIneath addresses: traditional BI vendors (Tableau, Power BI, Looker) that produce dashboards but not narratives; internal analytics teams that hand-author weekly decks; and emerging AI-first analytics tools (e.g., Hex Magic, Mode AI assistants, ThoughtSpot Sage) that focus on ad-hoc question-answering rather than recurring, scheduled, leadership-grade report authoring. | |||||
| What is the projected growth rate of your target market segment over the next 3-5 years? | The North American warehouse and DC operations software market is projected to grow at roughly 10–14% CAGR through 2028, driven by automation investment and the operational complexity of fulfilling omnichannel demand from the same physical nodes. The narrower category BIneath sits in — AI-assisted operations reporting and intelligence — is earlier in its curve and growing faster than the broader market, though forecasts vary because the category is still consolidating out of adjacent ones (BI, observability, conversational analytics). The directionally relevant point: leadership-facing AI reporting tools are projected to expand materially through the capstone timeframe and beyond, with adoption accelerating as enterprises modernize the underlying warehouse management systems that feed these reports. | |||||
| Business Model | What growth stage is your business currently in (e.g., startup, scale-up, mature)? | Mature business, mid-modernization on operational infrastructure. A Fortune 50 omnichannel mass retailer with a national DC network, decades of operational history, established leadership reporting cadences, and a stable enterprise data and tooling environment at the corporate layer. The specific DC in scope is a legacy site — roughly 30 years old, with physical and technological infrastructure behind the curve. The site is actively being modernized, but during the transition operational data lives in two worlds (legacy systems and the emerging modern WMS), and reconciling those two worlds for daily and weekly reporting is itself a meaningful share of the ICQA Manager's job. WHY NOW: this DC is also transitioning from a two-ICQA-Manager-per-site model to a single-ICQA-Manager solo-coverage model. This DC is piloting a future staffing model for the entire network. The work that was previously shared between two managers is consolidating onto one. BIneath is not a labor-displacement story — the staffing decision is independent of this product — but it is the operational answer to a staffing model that has already changed. Going from a 1-2 hour edit-required reporting cycle to a 5-minute review cycle is the difference between a workable solo role and one that absorbs the ICQA Manager's day in reconciliation and authoring at the expense of the floor work the role actually exists to do. BIneath is being built into this transitional substrate as an intentionally observational AI-native layer. This is appropriate for a supply chain that wants to become AI-forward but is not yet, and where employees across all levels need time to build trust with AI-assisted reporting before any prescriptive surface would be appropriate. | ||||
| How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?) | The business sells consumer goods at retail (apparel, home, hardlines, food, beauty) across physical stores and digital channels. Primary revenue model is retail transactional revenue, supplemented by owned-brand margin, retail media, and a small subscription tier. BIneath is an internal operations tool, not a revenue-generating product. Its business value is realized indirectly through: (1) ICQA Manager hours redirected from reconciliation and report authoring to floor work and team coaching, (2) faster and more consistent leadership decision cycles from higher-quality DBR/WBR inputs, and (3) reduced downstream cost from quality issues caught and addressed earlier. | |||||
| Who is your primary customer base (B2B, B2C, B2B2C)? | B2C for the parent business. BIneath itself is a B2E (business-to-employee) internal product, with the ICQA Manager as the primary user and OMs, SOMs, Site Director, and Operations Director as additional consumers of the reports BIneath produces. | |||||
| Differentiators | What are the key differentiators for your company? | Key differentiators of the parent retailer: integrated owned-and-operated supply chain (the business operates its own DC network rather than relying primarily on 3PLs), a strong owned-brand portfolio, an omnichannel fulfillment model in which the same DCs serve both store replenishment and digital orders, and a culture of operational rigor at site level that has produced rich, standardized internal data on quality and throughput. The active WMS modernization is itself a differentiator. Peer retailers are at varying points along this curve, and a site that brings AI-native reporting online during modernization will be ahead of the industry baseline once the modernization completes. | ||||
| Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.) | Customers | Who are the customers (ie buyers) of your product? | Internal product. The economic buyer is the DC operations leadership chain — Site Director and Operations Director — who sponsor internal tooling investments that reduce reporting burden on ICQA Managers and improve the consistency of leadership reviews. Regional VP and above are out of scope for the capstone. At that level, leadership consumes separate regional reporting at a different grain that BIneath does not produce. Regional-level reporting (Regional DBR/WBR/MBR/QBR authoring) is recognized as a natural future expansion for BIneath but is explicitly not in capstone scope. | |||
| End Users | Who are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context? | PRIMARY USER: ICQA Manager (Inventory Control & Quality Assurance Manager) at a single Distribution Center. Owns site-level quality reporting, authors the Daily Business Review (DBR) and Weekly Business Review (WBR), reconciles operational data across legacy and modern systems, and manages a team of 30–40 ICQA team members. The site is consolidating from two ICQA Managers to one (solo coverage); the primary user is the single ICQA Manager absorbing the full site's reporting load. The ICQA Manager is the person whose time and judgment BIneath is built to amplify. SECONDARY USERS: Operations Managers (OMs) and Senior Operations Managers (SOMs). • OMs each oversee 100–200 team members in each department across four operational keys (shifts). They cannot always make 1:1 meetings with the ICQA Manager and benefit from on-demand, self-serve access to current DBR/WBR output so they can review site-level quality data on their own schedule. The solo-coverage staffing model makes self-serve access materially more important — there are fewer hours in the ICQA Manager's week to field synchronous 1:1s. • There are four SOMs at the site, collectively overseeing 24 salaried OMs. SOMs benefit from not being blindsided in meetings when reports run late. This is a frequent occurrence when the ICQA Manager is buried in reconciliation work and the deck arrives at the meeting start time rather than well in advance. TERTIARY USERS: Site Director and Operations Director. They are the leadership audience for the WBR and the executive consumers of BIneath output. They do not author or edit, but the consistency and depth of the reports they receive directly shapes their operating decisions. OUT OF SCOPE: Regional VP and above. Regional leadership has its own reporting at a different grain and is not a user of site-level BIneath output. Recognized as a future expansion territory. Revenue impact: none of these roles are directly revenue-generating, but their collective decisions drive site-level quality outcomes that flow through to inventory write-offs, customer order cancellations, and labor productivity. Reducing the time the ICQA Manager spends authoring reports redirects that time into the floor work and team coaching that actually moves those numbers. Improving the consistency and depth of the reports OMs, SOMs, and tertiary leadership consume sharpens the operating decisions made at every level above. | ||||
| Current Products / Services | If you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers? | The current product is BIneath in its rules-based, deterministic form (the SDR Platform). It is the operational substrate the AI-native capstone version is being built on. The current platform has three main surfaces: 1. SDR Dashboard — multi-tab operational dashboard tracking Defect Rate trends across the four scopes (Total Building, Inbound, Breakpack, Packing), with deep dives by key, team member, zone, workstation, reason code, and freight type. Includes a Prior Day view and an Insights tab with automated anomaly callouts. Used daily by the ICQA Manager and accessible to OMs and SOMs. 2. Daily Brief and Weekly Brief — print-optimized one-page reports rendered from the same data. The Daily Brief surfaces Day 1 vs Day 2 deltas, the PNO and VCSNC watch lists, cancel-mix donuts, and workstation/zone concentration. The Weekly Brief surfaces WoW deltas, 8-week spike strips, fiscal-calendar scorecards (MTD/QTD/YTD), and top movers. 3. WBR Deck — a presentation-grade PowerPoint generator that produces the Monday WBR from the same data. Always renders a Title slide plus a Deep-Dive slide for the worst-gap path, with an optional Overflow slide for a second exceeding path. All narrative on these slides is currently template-driven and is enforced as observational (never prescriptive) by a forbidden-phrase guard in code. BEFORE THIS PLATFORM EXISTED: report authoring was an all-day Sunday-into-Monday affair. Quality data for DBR/WBR had to be pulled from Greenfield cards, Splunk dashboards, Apollo, and the legacy 3D WMS. Reconciled via brittle homegrown macros, manual pivot tables, and SharePoint workbooks where OMs hand-key volume counts. The current rules-based BIneath compresses that all-day workflow to roughly 15 minutes IF no edits are needed after the ICQA Manager uploads CSVs. WHAT THE DETERMINISTIC LAYER CAN AND CANNOT DO: rules-based BIneath can compute and display WoW and DoD relationships, but only along axes that have been explicitly coded. The daily and weekly outputs capture cancel mix and freight type mix because rules exist for them; receive-to-cancel age was added as a view because a rule was written for it. If a new correlation worth surfacing emerged in the data — for example, a pull-to-cancel age pattern in Breakpack — the deterministic layer would not raise it on its own. A new rule would have to be coded, deployed, and maintained for that relationship to appear in future reports. This is precisely where the AI-native version of BIneath unlocks the next leap: surfacing correlations across operational dimensions without each one having to be pre-specified in code, and converting edit-required cycles into a 5-minute review cycle (see Pain Points). | ||||
| User Value Map | Target Persona | Who is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc) | Internal user. Primary persona: the ICQA Manager (Inventory Control & Quality Assurance Manager) at a single Distribution Center. Profile: 5–15 years of DC and quality assurance experience, accountable for site-level Defect Rate against goal across three process paths (Inbound, Breakpack, Packing) plus the rolled-up Building rate, plus daily and weekly quality reporting to leadership. Manages a team of 30–40 ICQA team members. Authors the Daily Business Review every morning and the Weekly Business Review every Monday. Is not an analyst by title but operates analytically. Reconciling data across legacy and modern systems is a daily reality given the site's mid-modernization state. Time pressure is the defining constraint. The ICQA Manager's job is the floor and the team, not the laptop. Every hour spent on reconciliation and report authoring is an hour not spent on the operational work that actually moves the quality numbers being reported. This is the primary pain BIneath addresses. Multi-stakeholder consumer picture (also relevant for persona framing): OMs (100–200 TMs each, four operational keys) and SOMs (four total, overseeing 24 salaried OMs) consume the ICQA Manager's output. They have their own time constraints and cannot always meet synchronously with the ICQA Manager. Site Director and Operations Director consume the WBR output as their primary site-level quality input. | |||
| Journey Map (current-state) | What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product? | CURRENT STATE — Weekly Business Review (Monday morning, with rules-based BIneath in place): 1. ICQA Manager pulls weekly defect and demand CSVs from the source systems (modern WMS where available; legacy systems for the data not yet migrated). 2. ICQA Manager reconciles any inconsistencies between the two worlds, a step that varies in duration depending on how clean that week's data exports are. 3. ICQA Manager uploads CSVs to BIneath. Rules-based BIneath produces a Title slide and a Deep-Dive slide for the deterministic worst-gap path, with templated headlines. 4. ICQA Manager reviews the generated deck. • If no edits are needed: total time to publish is roughly 15 minutes from start to leadership delivery. • If edits are needed: total time expands up to 2 hours. The additional minutes are spent doing analytical work the deterministic tool did not do; parsing data manually to identify departments worth surfacing that didn't top the gap list, drawing new or expanded conclusions about long-running tolerance issues or outsized anomalies, and rewriting headlines and narrative prose to reflect those expanded conclusions while scrubbing for prescriptive language. 5. ICQA Manager walks the deck Monday at WBR with Site Director and Operations Director. CURRENT STATE — Daily Business Review (every morning, with rules-based BIneath in place): 1. ICQA Manager pulls or reviews daily defect and demand data. 2. Reconciles legacy/modern data as needed. 3. Uploads daily CSVs to BIneath. The Daily Brief renders as an in-app view, a downloadable PNG, and a PDF. There is no DBR slide deck; the Daily Brief is the artifact referenced during the DBR. 4. ICQA Manager identifies the top one or two things worth attention from the prior day and prepares what to say about them at the DBR. 5. ICQA Manager presents at the DBR. OMs and SOMs consume the Daily Brief here. HISTORICAL CONTEXT (pre-BIneath, for reference): before the current rules-based BIneath existed, this workflow was an all-day affair every Sunday into Monday. Data pulled by hand from Greenfield cards, Splunk, Apollo, and the legacy 3D WMS; reconciled via homegrown macros that broke easily; pivot tables and SharePoint workbooks with hand-counted volume. The current rules-based BIneath has already collapsed that into a 15-minute workflow when no edits are needed. The AI-native version targets a 5-minute review cycle consistently, including the cases that currently require edits. | ||||
| Pain-points | Where does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe? | Most severe and frequent pain points in the current-state journey (with rules-based BIneath in place): 1. The 'if no edits' condition fails often enough to matter. Rules-based BIneath deterministically picks the worst-gap path as the deep-dive focus. In a meaningful share of weeks (and days), the worst-gap pick captures the right story. In the remaining cases, the ICQA Manager must do analytical work the deterministic tool cannot — parsing data manually to identify a long-running out-of-tolerance pattern, an outsized anomaly that compressed in a single shift, or a department that has been drifting quietly for weeks without exceeding the worst-gap threshold. Drawing those expanded conclusions and rewriting the narrative to reflect them adds up to 2 hours per edit-required cycle and is the single largest variable cost in the workflow. Frequent enough, and high enough impact, to be the headline pain, and materially worse under the solo-coverage staffing model the network is moving to. 2. CSV upload as a friction point. The current platform requires the ICQA Manager to extract CSVs from source systems and upload them. Even when the rest of the workflow runs cleanly, the CSV step adds 5–10 minutes per cycle and is sensitive to source-system availability and format drift. Frequent (daily for DBR, weekly for WBR), moderate severity. 3. Reconciliation between legacy and modern data sources. The DC's mid-modernization state means quality data lives in two worlds, and the ICQA Manager is the human reconciliation layer. This pain varies in severity by week depending on which data sources are exporting cleanly. Frequent, moderate-to-high severity. Note: this pain is partly addressable through better data engineering and is mostly out of scope for this capstone's AI layer, but the AI layer's ability to query enterprise data directly (rather than depending on CSV uploads) is a step toward reducing this friction. 4. Narrative consistency varies with the author's day. Tired ICQA Manager, distracted ICQA Manager, ICQA Manager in the middle of a peak event. The narrative quality of the DBR/WBR fluctuates accordingly when edits are needed. Leadership receives reports that are uneven in voice and depth. Frequent on edit-required cycles, moderate-to-high severity. 5. Pattern recognition across weeks and days is human-held. The ICQA Manager authors each report largely from prior-week or prior-day data, holding multi-week trend context in Sharepoint archives. Identifying whether this week's Inbound gap is a return of a three-weeks-prior issue or is genuinely new requires human memory and interpretation the tool does not currently support. Frequent, moderate severity for ICQA Manager authoring; high severity for leadership consumption (leadership cares about trends, not weeks in isolation). 6. Consumer-side pain: OMs cannot self-serve current reports between 1:1s. OMs oversee 100–200 team members in each department across four keys and cannot always make synchronous meetings with the ICQA Manager. The solo-coverage staffing model makes this worse. Fewer hours in the ICQA Manager's week to field 1:1s; more reliance on the report itself as the communication artifact. When OMs want to review current quality data for their span, the report is either stale or not yet authored. Moderate frequency, moderate-to-high severity. 7. Consumer-side pain: SOMs get blindsided in when reports run late. When the ICQA Manager is buried in reconciliation work, the WBR can arrive at the meeting start time rather than ahead of it. SOMs walk into reviews without preparation time. Moderate frequency, moderate-to-high severity (reflects on SOMs and erodes trust in the reporting cadence). | ||||
| AI Opportunities | From your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first. | Pain points from F20 grouped by capstone scope, ranked most severe first within each group. ADDRESSABLE BY GENERATIVE AI, IN CAPSTONE SCOPE: Pain #1 (the headline pain) — Edit-required cycles. An LLM-based agent with reasoning over multi-week operational context can surface secondary callouts on the deck that the deterministic rules-based version cannot. Long-running or cyclical out-of-tolerance patterns, outsized single-shift anomalies, departments drifting quietly below the worst-gap threshold. The worst-gap path remains the deterministic spine of the deep-dive slide (preserving the trust gradient appropriate for a supply chain that is not yet AI-native). The AI layer adds secondary callouts that enrich the narrative with judgment the deterministic version cannot do. This converts the 1-2 hour edit-required cycle into a 5-minute review cycle on most weeks, a material unlock under the solo-coverage staffing model. Pain #4 — Narrative consistency. A scheduled agent producing narratives under a stable system prompt and a forbidden-phrase guardrail produces consistent voice across reports, independent of the author's day. The forbidden-phrase guard already exists in the codebase as a deterministic check; LLM-generated narrative is gated by it before any report is delivered. Pain #5 — Cross-week pattern recognition. The agent's planning loop has access to historical context (prior 8 weeks) as tool returns. It can identify whether this week's gap is continuing, returning, or new, and surface that context in the narrative without requiring the ICQA Manager to hold it in their head. Pains #6 and #7 — Consumer-side self-serve and on-time delivery. The scheduled autonomous run (Monday 9 AM Central for WBR) ensures the report is available before consumers need it. OMs and SOMs gain self-serve access to the most recent generated WBR. No ICQA Manager keystroke is required for the report to exist. ADDRESSABLE BY GENERATIVE AI, NOT IN CAPSTONE SCOPE: Pain #2 — Direct enterprise data querying. An agent with tool access to API endpoints can pull the data the report needs directly from the enterprise data layer, eliminating the CSV upload step. The agent's tool catalog is the existing OpenAPI contract, designed to be substrate-agnostic. In a future production deployment, the same tools could run against enterprise data warehouse endpoints rather than CSV-fed Postgres. In the capstone, the agent reads from the same CSV-fed Postgres the deterministic version reads from; replacing the CSV step is out of scope. Pain #3 — Legacy/modern reconciliation. The DC's mid-modernization state means quality data lives in two worlds, and the ICQA Manager is the human reconciliation layer. This is largely a data engineering problem rather than an AI problem, but the direct enterprise data querying described above would address one of its symptoms in a future deployment. | ||||
| Develop an AI Solution Hypothesis | AI Solution Hypothesis | Diverge | Ideate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage. | Solutions ideated against the high-priority pain points: A. Templated narrative generator — replace the current string-template headlines with LLM-generated prose seeded with the same structured inputs. Lowest-complexity, lowest-leverage. B. Conversational analyst — chat interface where the ICQA Manager asks questions and an LLM-backed system produces charts and answers from the operational data. C. Manager + Specialists Multi-Agent Architecture — a Manager agent reads the scorecard and dispatches process-path specialists (IB analyst, BPK analyst, PPS analyst) and a cross-path Movers analyst. Each specialist owns its domain prompt and a narrower tool subset, returning structured findings to the Manager. The Manager synthesizes findings and a separate Composer agent renders final callouts and speaker's notes under the forbidden-phrase guard. Same WBR output as Solution E, different decomposition. D. Single-purpose DBR Author Agent — same pattern as C, but for the daily review. Out of Phase 1 scope: there is no current-state DBR slideshow today, so building one would mean designing a net-new artifact rather than configuring the existing WBR pipeline for a different period. Recognized as a near-term Phase 2 expansion once a DBR slideshow artifact is established. E. Investigation Agent with Scheduled Invocation — a single planning agent that investigates the week's operational data, decides which secondary callouts and speaker's notes to surface alongside the deterministic worst-gap focus, and authors the AI layer of the WBR deck on a scheduled run. Architecturally extensible to additional periods (DBR) as future configurations. (The selected capstone solution.) F. Variance Investigator — interactive agent that responds to ICQA Manager questions like 'why did Inbound miss this week' by running a real multi-step investigation across operational panels. (Future v2 surface.) G. Autonomous Daily Watcher — agent that monitors continuously, decides when something is worth flagging, and notifies the ICQA Manager. (not scoped) H. Prescriptive Recommendations Agent — agent that suggests operational actions to leadership based on observed patterns. Distinct product surface from observational reporting. (Future Phase 2 — appropriate only after observational core has earned trust.) I. Multi-site Comparative Agent — agent that benchmarks one DC's metrics against peer sites and authors comparative narrative. (Future expansion territory.) J. Regional Reporting Agent — extension of the Report Author Agent to author Regional-level DBR/WBR/MBR/QBR for RVP and above. Same architecture, broader scope. (Future expansion, not in capstone.) K. Forecasting Agent — agent that projects forward (next week's expected rate, end-of-quarter trajectory) based on operational signals. (Future expansion.) | ||
| Converge | Rank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project. | Ranked by impact × feasibility within the capstone horizon: 1. (E) Investigation Agent with Scheduled Invocation. Highest impact: addresses the largest pains (edit-required cycles, narrative consistency, cross-period reasoning, consumer-side self-serve and on-time delivery) on the weekly reporting cadence. Feasible: shares one planning loop, one tool catalog, and one evaluation harness, with the architecture extensible to additional periods as future configurations. 2. (C) Manager + Specialists Multi-Agent Architecture — same WBR output as E but decomposes the investigation across domain specialists. Stronger pattern-match to the Manager Role agentic primitive, and specialist prompts can encode deeper per-path expertise. Costs against E: roughly 4x the prompt iteration work (Manager + three path specialists + Composer), roughly 4x the API cost per run, and an inter-agent coordination layer that doesn't exist in the single-agent version. Premature decomposition for Phase 1: the specialists do not perform meaningfully different kinds of work, only different domains, which is a split by domain rather than by capability. Higher feasibility risk in a 3-week capstone horizon. 3. (F) Variance Investigator — high impact for ICQA Manager ad-hoc analysis but interactive scope, sparser ground truth, and weaker evaluation signal. Best fit for a v2 surface after the scheduled report agent has earned trust. SELECTED FOR CAPSTONE: Solution E — the Investigation Agent with Scheduled Invocation, shipping with the WBR configuration at launch. The agent investigates the week's operational data through a typed tool catalog, produces secondary callouts that merge seamlessly into the deck's existing callouts list, and authors speaker's notes embedded in the PPTX notes pane. The DBR configuration is deferred to Phase 2 because there is no current-state DBR slideshow to extend. Adding one would mean designing a net-new artifact rather than configuring the existing WBR pipeline for a different period. ARCHITECTURE CHOICE: Solution E uses a single planning agent rather than the manager-with-specialists pattern (Solution C). The work decomposition by process path (IB, BPK, PPS) is by domain rather than by capability — all three paths require the same kind of reasoning over the same panel structure — which means a specialist split is premature decomposition for Phase 1. The capstone instead invests architectural budget in evaluation rigor, prompt iteration depth, and the layered guardrail stack, all of which yield more measurable quality improvements than agent multiplication in a 3-week horizon. The two-phase investigate/compose seam inside the single agent is deliberately type-level rather than network-level, which preserves the option to refactor to dedicated sub-agents in a future phase without restructuring. ARCHITECTURAL POSITION: the agent does not replace the deterministic worst-gap logic that drives the deep-dive slide today. The worst-gap path remains the spine of the deep-dive. The agent's value is in (a) authoring narrative prose under the existing forbidden-phrase guard, and (b) reasoning across multi-week context to surface secondary callouts and speaker's notes that enrich the deck: long-running tolerance issues, outsized anomalies, departments drifting quietly. The agent's callouts merge seamlessly into the existing callouts list with no visual differentiation from deterministic ones. Provenance lives in the persisted data layer for audit and debugging, not in the rendering. This is a deliberate product choice: the deck is the artifact, and AI authorship is a quiet capability rather than a labeled feature. The trust burden this places on output quality is borne by the evaluation suite (see Develop phase). Earning trust at the secondary-callout level, gated by rigorous evaluation, is the appropriate Phase 1 posture for a supply chain that is not yet AI-native. AUTONOMY POSTURE: scheduled runs on a weekly cron (Monday 9 AM Central, America/Chicago), with on-demand runs available via an internal route. The agent's callouts merge into the deck automatically — Level 4 autonomy for the merge step, conditional on a multi-layer evaluation and guardrail stack — while an advisory veracity layer (deterministic grounding plus an LLM judge) now surfaces per-callout flags in the deck ReviewPanel so the ICQA Manager reviews them before approving the week. The agent's output is gated by: (1) the forbidden-phrase guard applied at three points (composition, persistence, and merge into the deck), (2) a confidence floor at composition that drops observations below 30% rather than surfacing them with a low number, (3) a bare-code regex guard that rejects any unlabeled reason code (R23, I29, PNO, etc.) before it can reach the slide, (4) graceful degradation when guardrails fail. The agent's output for that period is suppressed and the deck publishes with deterministic content only, with an internal alert raised for inspection. (5) a golden-set evaluation pass requirement as a precondition before any prompt iteration ships to production. This posture matches the existing deterministic publication pattern (data ingest auto-generates briefs) and is appropriate only because the eval architecture supports it. PRODUCT VISION BEYOND CAPSTONE: BIneath is framed as a multi-agent operations intelligence platform. The Investigation Agent (this capstone, in its WBR configuration) is the first agent. The near-term Phase 2 extension is the DBR configuration of the same Investigation Agent, which becomes feasible once a DBR slideshow artifact exists for the agent to author into. A future Variance Investigator (solution F) becomes a second product surface for interactive analysis. A future Prescriptive Companion Mode (solution H) layers on top of the observational core as a distinct, opt-in surface with its own guardrails, its own evaluation rubric, and its own human-in-the-loop confirmation pattern. This is appropriate only after the observational agent has demonstrated reliability and built trust across the multi-stakeholder consumer base. Regional reporting (solution J), multi-site comparison (solution I), and forecasting (solution K) are recognized as natural extensions of the same agent architecture but are explicitly not prioritized for the capstone or near-term roadmap. | ||||
| DESIGN | Define Target State Workflow | UX Flows & Wireframes Suggested Tool: Excalidraw | Workflow (future) | Assuming your product or feature works as desired, what is the target state workflow? | TARGET STATE — Weekly Business Review (AI-native BIneath, now built and published to production): 1. ICQA Manager uploads the week's CSVs to BIneath, as in the current state. The CSV upload step is unchanged in Phase 1; the AI layer reads from the same CSV-fed Postgres the deterministic version reads from. 2. The weekly cron fires Monday 9 AM Central (America/Chicago), calling runWeeklyCronTick, and is idempotent (skips a week already generated). Reliable weekly firing assumes an always-on (Reserved VM) deployment; see Technical Readiness. The same run can also be triggered on demand via an internal, X-Internal-Token-gated HTTP route. 3. Investigation phase: the agent reads the deterministic deck's structured payload to understand what is already on the slide, then plans a deeper investigation across the full operational dataset (every reason code, every workstation, every team member, every freight category) using investigation-tier API endpoints. 4. Composition phase: the agent selects the strongest candidate observations, writes secondary callouts that merge seamlessly into the deck's existing callouts list, and authors speaker's notes embedded in the PPTX notes pane. Every claim carries a confidence percentage and citation in the persisted payload. 5. Guardrail pass: the narrative guard (forbidden-phrase, bare-code regex, reversal-marker), the confidence floor, and the goal-direction check all run on the agent's output. Failures trigger up to two retries before the offending claim is dropped. 6. Advisory veracity pass: deterministic grounding plus a best-effort LLM judge run on the composed callouts and notes and persist per-callout findings (keyed by surface/slot/index) to agent_outputs.veracity. This is advisory only — it never blocks save or approve, grounding always runs first, and a judge failure still returns grounded claims. 7. The merged deck publishes. OMs and SOMs access it through the existing self-serve interface; Site Director and Operations Director receive it ahead of the WBR meeting. 8. Review and approve: in the deck ReviewPanel the ICQA Manager sees each callout and note with its advisory veracity badges (grounding findings + judge verdicts), can edit or drop any item, and approves the week. Status flows draft → edited → approved on agent_outputs. In most weeks this is a roughly 5-minute confirmation; in the small share of weeks where edits are warranted, the ICQA Manager edits in place. Like/dislike + reason-tag + note feedback (and the implicit draft-vs-approved diff) is distilled into a bounded lessons block that steers the NEXT draft only; it never alters an already-served deck. 9. ICQA Manager walks the deck at the WBR review with Site Director and Operations Director. WHAT CHANGED FROM CURRENT STATE: • The deterministic worst-gap deep-dive is unchanged. The agent's editorial authority is bounded to secondary callouts and speaker's notes. • Multi-week pattern recognition is now in the agent's planning loop rather than in the ICQA Manager's head and SharePoint archives. • Reports publish on a predictable cadence without ICQA Manager involvement. The SOM blindside pain (Pain #7) is structurally eliminated. • An advisory veracity layer now surfaces grounding + judge flags per callout in the ReviewPanel before approval, and a reviewer learning loop folds feedback into the next draft. • The 1–2 hour edit-required cycle collapses to a roughly 5-minute review cycle in most weeks. The ICQA Manager's time is reclaimed for floor work and team coaching, which is materially important under the solo-coverage staffing model. WHAT DID NOT CHANGE: • The deck's visual structure (Title slide, Deep-Dive slide for worst-gap path, optional Overflow slide for a second exceeding path). • The deterministic worst-gap logic that drives the deep-dive selection. • The narrative guard as the observational-not-prescriptive boundary. • The ICQA Manager's role as the human reviewer of record before WBR presentation, now supported by per-callout veracity badges. • The CSV upload step. The ICQA Manager still extracts CSVs from source systems and uploads them. The agent reads from the resulting Postgres, same as the deterministic version. Replacing the CSV step with direct enterprise data querying remains a Phase 2 architectural opportunity. | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Build Wireframes | Wireframes | How will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features? | https://drive.google.com/file/d/1CqkcxtuxZSP438ejWoCSrsYpCH_p2ax_/view?usp=sharing WBR Investigation Agent — system walkthrough The agent runs on a weekly cron, not by user trigger. Every Monday morning the cron fires runWeeklyCronTick, which calls POST /api/agent/run-weekly (an internal endpoint gated by an X-Internal-Token header). That endpoint kicks off a two-phase Claude Opus 4.7 conversation. Phase one is the investigator: it's given the deterministic deck payload for the week as context, then loops through tool calls against /api/investigate/* endpoints — un-truncated reads of reason codes, freight categories, team-member movers, workstation Paretos — building up a structured list of CandidateObservations with confidence scores and citations. Phase two is the composer: it takes those candidates and writes the final slide-face callouts and speaker notes, with every claim passing through the narrative-guard's forbidden-phrase check and bare-code regex (up to two retries before a claim is dropped). The composer's output lands in a new agent_outputs table, keyed by weekEnd. From there, every /api/wbr-deck request runs mergeAgentIntoDeck server-side: it reads the agent row for the requested week, runs the guard again as defense-in-depth, applies per-slide caps (title 0-or-2, deep-dive 0/2/4/6, overflow 0/2/4/6), and routes each agent callout to the correct slide based on its pp tag. The merged payload flows out to two render surfaces — the rendered deck (title, deep-dive, optional overflow) and the PPTX speaker-notes pane when the user clicks Export PPTX. Title slide. The slide has the same shape it always had: scorecard cards for Building, IB, BPK, PPS along the top; four 8-week trend mini-charts below; a callouts list at the bottom. The load-bearing product decision lives in that callouts list. Deterministic callouts (worst-gap headline, weeks-above-goal counter) and agent callouts (secondary patterns, individual-level signals) sit in the same list, in identical typography. The audience reads them as one set of observations. No badges, no "AI-generated" labels, no source attribution on the slide face. The mergeAgentIntoDeck function enforces an even-count cap on title callouts — either zero agent callouts or exactly two — so the slide never feels lopsided. Deep-dive slide. When at least one path misses goal, the deep-dive slide shows the worst-gap PP. The layout has a header KPI block (rate, goal, vs-goal pp, solds, cancels), a by-key table breaking the path across A1/A2/B1/B2 shifts, a top-movers panel of TM-level WoW changes, and path-specific panels — reasons and freight mix for IB, workstation/zone Paretos and PNO watch for BPK, zone Paretos and VCSNC watch and BPK pullers for PPS. The agent never writes into the structured panels. Its only seat on this slide is the callouts list, same merge-invisibly rule as the title. Items route to this list when their pp tag matches the deep-dive's PP, sorted by confidence, capped at six. Overflow slide. When two paths miss goal and the second one is far enough above goal to warrant attention (rate-to-goal ratio ≥ 1.5), an overflow slide renders for the second-worst PP. Same layout as the deep-dive. Agent callouts tagged with that PP route here instead of to the primary deep-dive. Speaker-notes pane. This is the one place where agent-generated content has a surface of its own. After the user exports the PPTX, opening the file in PowerPoint shows the deterministic slide canvas above and an agent-written paragraph below in the notes pane — one paragraph per slide (title, deep-dive, overflow), denser than the slide-face callouts, written in the voice of a presenter who's looked at the data. Confidence scores and citations persist in agent_outputs for the eval harness but are stripped before the prose reaches the notes pane. Failure and empty states. Three graceful-degradation paths matter. (a) Agent hasn't run yet: getAgentOutput(weekEnd) returns null, the deck renders with deterministic callouts only, the PPTX export ships with no speaker notes. Nothing breaks; the deck looks like its pre-agent self. (b) Agent ran but every callout was dropped by the narrative guard: same render as (a). Dropped observations still persist in agent_outputs for later inspection by the eval harness, but nothing reaches the deck. (c) All paths within tolerance: the deep-dive slide is replaced by a tolerance slide ("All paths within tolerance"). Agent still emits title-level callouts and speaker notes if its confidence holds; no deep-dive callouts route because there's no deep-dive slide to route them to. Review and veracity surface (ReviewPanel). Before approval, the deck ReviewPanel renders each agent callout and speaker note with advisory veracity badges: deterministic grounding findings plus best-effort LLM-judge verdicts, keyed by (surface, slot, index) so each flag sits on the exact item it describes. The reviewer can edit or drop any item, leaves structured like/dislike + reason-tag + note feedback, then approves; status moves draft → edited → approved on agent_outputs. The veracity check is advisory only: it never blocks save or approve, grounding always runs first, and a judge failure still returns grounded claims. Feedback is distilled into a bounded lessons block that steers the next draft and never alters an already-served deck. Production run model. The agent runs on an in-process node-cron tick (Monday 9 AM Central, America/Chicago), idempotent across re-runs, with on-demand generation and veracity recompute available via internal-token-gated HTTP routes. A reliable weekly cron requires an always-on (Reserved VM) deployment; autoscale scales to zero and the in-process timer will not fire dependably. | |||
| Develop Prototype to showcase AI interactions | Prototype Screens Suggested Tool: lovable.dev | What aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases? | https://drive.google.com/file/d/1FW_DFRYDLLBjPFmatq6Vvn8kblLIi3ns/view?usp=sharing The Phase 1 prototype is functional: BIneath in its AI-native form runs end-to-end against historical weeks, with the Investigation Agent producing real secondary callouts and speaker's notes that render through the existing pptxgenjs pipeline. The prototype is not a mockup; it is the working agent operating against the real codebase, now built and published to Replit production. WHAT THE PROTOTYPE DEMONSTRATES: 1. The agent investigating a historical week. The investigation phase reads the deterministic deck's payload, then calls investigation-tier API endpoints to look beyond the slide's top-N truncation (every reason code, every workstation, every team member, every freight category). The tool call sequence is visible in the persisted run record. 2. The agent composing. The composition phase takes the candidate observations from the investigation phase and produces secondary callouts (0 or 2 on the title slide, 0/2/4/6 on the deep-dive and overflow slides) and speaker's notes (4-6 per slide) under the forbidden-phrase guard and the bare-code regex guard. 3. The merged deck. The agent's secondary callouts merge into the existing deterministic callouts list with no visual differentiation. Speaker's notes appear in the PPTX notes pane in flowing prose. 4. The guardrail stack catching real failures. The forbidden-phrase guard, the bare-code regex, and the confidence floor each surface in the run logs when they fire. 5. The before/after contrast. Opening the same week's deck with and without the AI layer makes the agent's contribution visible: deterministic deck alone is templated and serviceable; deck with AI layer is closer to what an experienced presenter would actually say. 6. The advisory veracity layer and ReviewPanel. Deterministic grounding findings plus LLM-judge verdicts render as per-callout badges, and the reviewer edits, drops, or approves with draft → edited → approved status before the week is finalized. FEATURES ESSENTIAL FOR LAUNCH (Phase 1): • Investigation phase with investigation-tier tool access • Composition phase with secondary callouts (slide face) and speaker's notes (notes pane) • Confidence percentage and citation persisted with every observation • Forbidden-phrase guard, bare-code regex guard, confidence floor • Scheduled cron run on Monday 9 AM Central (America/Chicago), idempotent, plus on-demand generation and veracity recompute via internal routes (reliable weekly firing requires an always-on Reserved VM deployment) • Advisory veracity layer (deterministic grounding + best-effort LLM judge) surfaced per-callout in the deck ReviewPanel before approval • Review/approve flow with draft → edited → approved status, and a reviewer learning loop that distills like/dislike + reason-tag + note feedback into a bounded lessons block steering the next draft • Graceful degradation: agent output suppressed and deck publishes deterministic-only if eval gate fails FEATURES DEFERRED TO LATER RELEASES: • DBR configuration (Phase 2, contingent on a DBR slideshow artifact being established) • Variance Investigator interactive query surface (Phase 2) • Prescriptive Companion Mode (Phase 2 or later, with its own guardrails and confirmation pattern) • Regional reporting at MBR/QBR cadence (future expansion) • Multi-site comparative narrative (future expansion) • Forecasting (future expansion) Now built and published to Replit production: an Express API behind the path-routed proxy with a /api/healthz health check and a source-mapped production bundle, the weekly Central cron plus on-demand routes, and a production Postgres database separate from development populated via the Publish flow. It began as a functional prototype running end-to-end against historical weeks on the real codebase, not a mockup. | |||
| Initial Prompt Design | Master Prompt [Initial Design] | Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency? | The agent runs as a two-phase Claude Opus 4.7 conversation (max_tokens: 8192). The two phases share one Anthropic conversation context but use distinct system prompts, with an explicit type-level seam between them in code (the investigator's output is a structured CandidateObservation[] consumed by the composer). The seam is in code rather than across network calls in Phase 1, but the architecture supports a future refactor to dedicated sub-agents without restructuring. INVESTIGATOR PHASE — purpose: read the deterministic deck's payload, plan an investigation across the full operational dataset, return a structured list of candidate observations. Key prompt elements: • Domain semantics: explicit definition that defect rate, cancel counts, and reason-code counts are all minimized. A value below goal is favorable (meeting goal); above goal is unfavorable (missing goal). Confusing the direction is explicitly flagged as the most damaging error. • Goal language vocabulary: 'missing' (above goal), 'meeting' (at-or-below), 'at goal' (exactly), 'well under / comfortably below / far below' (clearly below). • WoW reasoning rule: the WBR reports on the most recently completed fiscal week. That week IS 'last week' from the WBR's perspective. Earlier comparisons must use 'Week N' explicitly or 'the week before last.' • Tool catalog: two tiers. Slide-visible tools return the same top-N truncated data the deterministic deck consumes (so the agent knows what is already on the slide). Data-deep tools return un-truncated operational data. • Investigation methodology: the agent looks for what the deterministic worst-gap logic does not surface. Long-running patterns, cross-dimensional observations the top-N truncation hides, composition shifts, material individual contributors with counterfactual impact. • Hard data rules: SVIAITM excluded from all team-member commentary; PPS team-member movers restricted to the VCSNC-only path (cancelUserId T560321); BPK Flow is a real freight bucket, not collapsed into N/A. • Team member references: real names when the roster resolves them; Z-logins with 'likely terminated' when it does not (the roster reflects only active employees at report time). • Output: structured JSON array of CandidateObservation objects with kind ('callout' or 'note'), preferredSlide, pp (process path tag for routing), claim, confidence (whole-number [30, 95]), citation, and supporting notes. COMPOSER PHASE — purpose: select the strongest candidates, produce final secondary callouts and speaker's notes in the voice of an experienced site quality manager. Key prompt elements: • Two output surfaces: speaker's notes (4-6 observations per slide, flowing prose in the PPTX notes pane) and secondary callouts (0, 2, 4, or 6 per slide, merged into the existing callouts list with no visual differentiation). • Voice calibration: one short example of target voice. Past tense for the week's events. Dense comma-spliced prose welcome. Named individuals when contribution is material, paired with counterfactual. Calibrated hedging ('closer to,' 'appears to,' 'roughly,' 'may be'). End paragraphs with synthesis when the data builds toward one. • Verbalization conventions: path abbreviations (IB, BPK, PPS, Total Building); key + path ordering (key first: 'A2 Packing,' never 'PPS A2'); period references ('Week 18' or 'Week ending May 9,' never 'fiscal'); arrow notation for transitions ('816 → 1,057'); tilde for approximations ('~60%'); rates as percentages; counts as '127 cancels' not 'cans'; demand in relative terms ('a third less demand') not absolute counts; reason codes always by human-readable label ('IB WIP') not by code ('R23'). • Goal-direction check: explicit worked example of CORRECT and INCORRECT phrasing for a hypothetical scorecard. The downstream guard rejects miss-language about below-goal performance, so the composer is told to check scorecard[pp].vsGoalPp before writing any sentence with 'miss' in it. • Guardrails: observational never prescriptive. Forbidden territory enumerated ('should,' 'needs to,' 'must,' 'recommend,' 'priority,' 'coaching,' etc.). Confidence floor 30 (below dropped); ceiling 98 for composer output. • Output: JSON object with secondaryCallouts (per-slide arrays) and speakerNotes (per-slide arrays). Confidence and citation persisted with each observation but not surfaced in the rendered deck. The verbatim text of both system prompts is maintained in the codebase at lib/agent-investigator/src/phases/investigate.ts and lib/agent-investigator/src/phases/compose.ts and is documented in the feature log. | |||
| Prepare for Testing & Iteration | Evaluation Criteria & Test Plan | Evaluation Criteria | What specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output? | The evaluation suite is load-bearing infrastructure rather than a checkpoint. The agent's callouts merge into the deck automatically (Level 4 autonomy for the merge step), but an advisory veracity layer — deterministic grounding plus a best-effort LLM judge — now surfaces per-callout flags in the deck ReviewPanel so the ICQA Manager sees them before approving the week. The suite is the trust mechanism that makes seamless callout merge and that advisory pre-approval review defensible. The veracity layer is advisory only: it never blocks save or approve, grounding always runs before the judge, and a judge failure still returns grounded claims. GOOD OUTPUT means an agent run that satisfies all four criteria categories below. Failure on any category is graded; some failure modes trigger graceful degradation (deck publishes deterministic-only and an alert is raised); others trigger composer retries (up to two). 1. FACTUAL ACCURACY (mechanically checkable): • Goal-direction correctness: every claim that uses 'missing,' 'meeting,' 'above goal,' 'below goal,' 'exceeded,' 'hit' must match the actual rate vs. goal direction in the scorecard payload. This is the rubric item that catches the IB miss-vs-meet error class. • Number matching: every numeric value in the agent's output must appear in the underlying tool returns. No invented numbers. • Path attribution: claims tagged with a specific process path (pp) must match the path whose data the claim is drawn from. • Reason-code resolution: any reference to a defect reason must use the human-readable label, never the bare code. The bare-code regex guard enforces this; the eval verifies the guard fired correctly where applicable. 2. SEMANTIC CORRECTNESS (LLM-as-judge or human-graded): • Cross-dimensional observation validity: when the agent says 'A2 had 2x defects despite a third less demand,' the underlying data must actually show that relationship. • Multi-week pattern claims: when the agent says 'BPK has been above tolerance for 4 consecutive weeks,' the 8-week tool return must support the claim. • Materiality: observations should be worth surfacing — observations a thoughtful presenter would actually make. Observations that are technically accurate but trivial (e.g., restating a number already on the slide) fail this category. • Citation traceability: every observation traces back to one or more tool returns that grounded it. 3. VOICE & VERBALIZATION (mechanical + LLM): • Forbidden-phrase compliance: no observation contains words from the forbidden list (should, needs to, must, recommend, priority, coaching, address, fix, improve, etc.). The narrative-guard enforces; the eval verifies. • Verbalization conventions: path abbreviations correct; key-first ordering ('A2 Packing,' not 'PPS A2'); period references use 'Week N' or 'Week ending,' never 'fiscal'; reason codes use labels; arrow notation for transitions; tilde for approximations. • Tone: observational, never prescriptive. Past tense for the week's events. Calibrated hedging when appropriate. 4. CALIBRATION: • Confidence in [30, 95] for investigator output, [30, 98] for composer output. Observations below 30 must be dropped. • Stated confidence should match the strength of the underlying evidence (multi-period support, large effect size, dense data → higher confidence; single period, small effect, sparse data → lower confidence). • Hedging language matches the confidence level. A 90% claim should not read as a tentative one; a 45% claim should read with appropriate caution. EVAL GATE FOR PRODUCTION RUNS: the eval architecture operates at two tiers. At the claim level, the narrative guard catches forbidden phrases, bare reason codes, and confidence violations as the composer produces output. Per-claim failures trigger up to two composer retries; if the claim still fails the guard, it is dropped from the run. At the run level, the eval gate requires that the surviving claims (after retries and drops) pass all factual-accuracy checks (mechanical) and at least 80% of semantic-correctness checks against the week's goldens. A run that fails the run-level eval gate triggers graceful degradation: the agent's output for that week is suppressed entirely, the deck publishes with deterministic content only, and an internal alert is raised for inspection. | ||
| Example Cases | What specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs? | TYPICAL CASE — Monday WBR for a week with one path missing goal: Input: weekly defect and demand data for a representative week. Scorecard shows Total Building 1.42% against 1.38% goal, IB 0.91% against 0.76% goal (missing by 0.15pp, the worst gap), BPK 2.40% against 2.57% goal (meeting), PPS 1.05% against 1.23% goal (meeting). 8-week trends available. Full operational dataset accessible (every reason code, workstation, team member, freight category, zone). Expected agent behavior: • Investigator reads the slide-visible payload to understand what the deterministic deck already shows (IB worst-gap deep-dive, IB reasons donut, IB freight donut, IB WIP method × category, etc.). • Investigator queries data-deep endpoints for what the slide does not show (every IB reason code beyond the top 3, every workstation beyond the top 5, the full team-member list, freight categories beyond the top 3). • Investigator surfaces 6-10 candidate observations covering: IB-specific second-order signals the worst-gap deep-dive does not capture (e.g., a workstation outside the top 5 with an outsized week, a freight category whose share shifted week-over-week); long-running patterns on non-worst-gap paths (e.g., BPK has been at-tolerance but trending up for three weeks); cross-dimensional observations (e.g., a key in BPK with 2x defects despite a third less demand); material individual contributors with counterfactual impact. • Composer selects the strongest candidates. Produces 2 secondary callouts on the Title slide highlighting second-order findings the headline does not cover; 4 secondary callouts on the IB Deep-Dive merged into the existing callouts list with no visual differentiation; speaker's notes (4-6 per slide) in flowing prose, past tense, voice calibrated to an experienced site quality manager. • Every observation carries a confidence percentage [30, 98] and a citation to the tool return that grounded it. Confidence and citation are persisted but not rendered on the slide. • The forbidden-phrase guard, the bare-code regex guard, and the confidence floor all run on the composer's output. Any failures trigger up to 2 composer retries before the offending claim is dropped. Expected output quality: • Factual accuracy: IB is described as 'missing' the goal (correct direction). Numbers cited match the underlying data. Reason codes appear as labels (e.g., 'IB WIP'), never as raw codes (e.g., 'R23'). • Semantic correctness: secondary callouts surface observations the deterministic worst-gap logic does not surface. Cross-dimensional claims (defects vs. demand share, anomalies vs. baseline) are valid against the underlying data. • Voice: past tense throughout. Dense prose in speaker's notes. Key-first ordering when keys are mentioned. Arrow notation for transitions. Hedging proportional to confidence. • Calibration: confidence levels match evidence strength. High-confidence observations read plainly; lower-confidence observations carry appropriate hedge. Pass/fail rationale: a passing run delivers a deck that the ICQA Manager can present at Monday WBR with at most a 5-minute review and no substantive edits. A failing run either (a) gets caught by the eval gate and triggers graceful degradation (deterministic-only deck publishes), or (b) the ICQA Manager catches a failure during the 5-minute review and edits before presentation. Additional case categories (edge cases, all-paths-meeting-goal tolerance case, multi-path-missing case) are covered in the evaluation set (see Develop phase, F39). | ||||
| DEVELOP | AI Model Selection & Justification | AI Model Selection & Justification | Which AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product? | The solution uses Claude Opus 4.7 (claude-opus-4-7) via the Anthropic Messages API through the Replit AI Integrations proxy. Opus was selected for three reasons the task demands: (1) multi-step tool-use reasoning over multi-week operational data (the investigator runs an autonomous tool loop, deciding what to pull next), (2) strong instruction-following for the strict verbalization conventions and non-prescriptive guardrails, and (3) adaptive thinking, a private reasoning channel that returns a clean JSON block, which the composer and judge rely on for parseable output. Both phases run at max_tokens 8192 (judge 4000). Limitations: per-run latency (a bounded ≤8-iteration tool loop) and cost, and non-determinism, all mitigated by the deterministic grounding gate, the narrative guard, and the advisory veracity layer rather than by trusting the model. Integration is a single shared Anthropic client reading the proxy base URL and key from the environment; no model keys live in application code. | Please leave this area blank. This space is for the Instructor to provide you with feedback. | |
| Define Inputs | Input Specification Table | Required Fields | What are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement. | The only externally supplied field is weekEnd, the fiscal week-ending date (format YYYY-MM-DD), provided by the weekly cron or an operator. All other inputs are derived server-side from CSV-fed Postgres and handed to the agent as ground truth: the deterministic deck payload (/api/wbr-deck — scorecard rate/goal/vsGoalPp/WoW per scope, 8-week trend, headlines, per-PP deep-dive panels) as the slide-visible tier, and the data-deep endpoints (/api/investigate/* — un-truncated reason codes, workstations, zones, freight categories, team-member movers) as the data-deep tier. Internal transport is JSON over HTTP gated by an X-Internal-Token header. Source tables: weekly_demand, defect_records, daily_*, team_members. | ||
| Optional Fields | Are there any optional or user-customizable fields? How do they impact the AI’s output? | • lessons (optional): a bounded block of distilled reviewer guidance injected into the prompts; steers what the investigator surfaces and how the composer words it, but never invents facts. CLI/eval callers omit it. • judge (on/off): toggles the LLM-judge layer for advisory veracity and live evaluation; off = fast grounding-only. • edited flag: recompute veracity against an edited surface instead of the raw draft. • Per-week overrides (WEEKLY_OVERRIDES, incl. vcsncLoginMerges): correct team-member name/key/callout issues and fold a TM's two Z-logins into one for the VCSNC watch list. These affect identity/labeling and movers, never the underlying rate math. | ||||
| Define Good Output | Output Evaluation Checklist | Objective Criteria | What criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance) | The deterministic grounding rubric. A “good” claim: (1) cites a well-formed, resolvable endpoint+fields; (2) every count it states matches the cited surface exactly, and every rate matches within ~0.07pp; (3) every curated label (reason/method/freight) appears in the cited source; (4) it trips none of the guard checks (no bare codes, no forbidden/prescriptive phrases, no self-correction markers). A claim passes when it carries no finding at or above the failure severity (high by default); counts, guard violations, and unresolved citations are high-severity. Golden good claims must pass clean, with zero findings of any severity. | ||
| Subjective Criteria | Are there any criteria that require human judgment or qualitative assessment? | Captured by the LLM-as-judge rubric plus human review. The judge scores each claim 0–5 on voice (declarative, past-tense WBR register), accuracy (every figure/attribution supported by ground truth), and guard (free of forbidden phrases, bare codes, self-corrections); ≥3 on every dimension to pass. Human judgment in the ReviewPanel adds relevance, executive-readiness, and tone, captured as like/dislike + reason-tag + note feedback. | ||||
| Prompt Design Iteration | Master Prompt [Final Design] | Prompt Version 1 | What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints. | Two distinct system prompts (verbatim in the project's system-prompts doc): • Investigator prompt — domain semantics (defect rate minimized; four scopes; goal-language vocabulary; the “last week” rule), two-tier tool instructions, hard data rules (exclude the SVIAITM automation account; PPS movers via the VCSNC-only path; never collapse “BPK Flow” into N/A), team-member/Z-login handling, and a strict CandidateObservation[] JSON output contract with a confidence floor of 30 and ceiling of 95. • Composer prompt — voice calibration (a dense site-quality-manager register, with one worked example), verbalization conventions (paths, keys, periods, numbers, labels), the hard constraints the downstream guard enforces, and a strict JSON output (secondaryCallouts + speakerNotes keyed by slide) with per-slide caps. | ||
| Prompt Iterations | If revised, what changes did you make and why? How do you track and record prompt evolution? | Tracked through the build history. Key iterations: split a single prompt into the two-phase investigate/compose design with a hard code boundary (composer never sees the investigator transcript) to stop self-correction leakage; raised the confidence ceiling 95 → 98; added a dedup guard so title callouts don't echo the headline; added per-slide callout caps; added a “cans” → “cancels” scrub, a reversal-marker guard, and a worked goal-language example after observing miss-direction and slang failures; synced the curated reason-label map into the prompt so labels reconcile against source. Reviewer feedback is folded in continuously as a bounded lessons block (Tone/Brevity/Phrasing → composer; Veracity/Accuracy/Relevance → investigator). | ||||
| Data Preparation & RAG Implementation | Data Preparation & RAG Implementation | What data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information? | No vector RAG / embeddings. The agent uses structured, citation-scoped tool-augmented retrieval over the operational Postgres: deterministic SQL-backed endpoints expose the scorecard/trend/deep-dive (slide-visible) and un-truncated dimensional breakdowns (data-deep); the investigator pulls only what it needs via tool calls, and every claim must cite the endpoint + fields it used. “Data preparation” is the CSV → Postgres seed plus the deterministic aggregations the deck already computes, with curated label maps normalizing raw warehouse codes (reason/method/freight) to human labels. Because grounding reconciles each claim against the exact cited surface (not the union of the week), retrieval correctness is enforced rather than approximated. | |||
| Create Evaluation Set | Example Input/Output Data for Testing | Typical Examples | What are the most common inputs and expected outputs? Use real data if possible. | Golden fixtures of a representative week's good claims, each carrying its citation and required to pass the grounding check clean. Typical case: a week where one path misses goal; expected output surfaces the worst-gap deep-dive plus a second-order pattern the deterministic deck cannot (a multi-week drift, a concentration outside the top-N, or a material individual contributor with a counterfactual), all grounded to cited fields. | ||
| Edge Cases & Negative Cases | What examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain) | Golden “bad” fixtures, each declaring the finding kind it must trip: unmatched-count (fabricated number), unmatched-rate, bare-code (e.g. R23/I29), forbidden-phrase (prescriptive language), reversal-marker (self-correction), label-not-in-source, and unresolved-citation (cited surface absent). Operational edge cases: missing/sparse week data, all paths meeting goal (no deep-dive), judge unavailable (degrade to grounding-only), and terminated contributors surfacing as Z-logins. | ||||
| Test Example Data & Review Results | Manual Review | Run your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why? | The deck ReviewPanel is the manual-review surface: a reviewer sees each callout/note with its advisory veracity badges (grounding findings + judge verdicts), can edit or drop it, and leaves structured feedback before approving. Status flows draft → edited → approved on agent_outputs. | |||
| Automated Evaluation | What pass/fail rate or scores did the AI achieve on core criteria? | The deterministic grounding gate runs as a CI/validation step (agent-evals) over the golden fixtures and exits non-zero on any failure; good fixtures must pass clean, bad fixtures must surface their declared findings. Live spot-checks add the LLM judge (voice/accuracy/guard, each ≥3). Per draft, the advisory veracity report records grounding pass/flag counts (by finding kind) plus judge verdicts. | ||||
| Handle Edge Cases & Iterate | Edge Case Identification | What edge cases did you identify in testing or real usage? | Identified in testing and real usage: bare reason codes leaking into prose (dropped by the guard; fixed via curated labels); miss-language about a below-goal path (guard drop + a worked example added to the prompt); the “cans” slang (scrubbed); the composer “continuing” the investigator's reasoning and leaking self-corrections (fixed by the hard phase boundary); a TM's two Z-logins splitting the VCSNC watch list (fixed via a login-merge override); callouts routed to the wrong process-path slide (fixed via pp-routing). | |||
| Updates & Adjustments | What prompt or system adjustments have you made based on failures, feedback, or edge case observations? | Each failure mode produced a concrete fix: a prompt rule, a guard rule, a scrub, a per-week override, or an architectural change (the investigate/compose boundary). Reviewer feedback distilled into the bounded lessons block continuously adjusts what is surfaced and how it is worded for the next draft, without ever altering the served deck. | ||||
| Automate Evaluation Approach | Evaluation Method | What is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets? | Hybrid: (1) a deterministic citation-scoped grounding script — the gate, no LLM, run on fixtures and on every draft as advisory veracity; (2) an LLM-as-judge model grader for the softer voice/accuracy/guard rubric (live + advisory); (3) human review in the ReviewPanel. It scales by adding golden fixtures (a week + its expected outcomes) and via a live mode that pulls the latest persisted output for any real week. | |||
| Evaluation Frequency | How often will you re-run evaluations for new data, new prompts, or post-launch monitoring? | Grounding fixtures run in CI on every change (the agent-evals validation step); advisory veracity (grounding + best-effort judge) runs at draft generation and on-demand before each weekly review; live judge spot-checks as needed. Post-launch, every weekly run is grounded and optionally judged before approval. | ||||
| DEPLOY | Finalize Launch & Rollout Plan | Operational Readiness Checklist | Technical Readiness | Is infra (APIs, databases, rate limits, monitoring, rollback) tested and documented? | Published to Replit as an Express API behind the path-routed proxy, with a /api/healthz startup health check and a source-mapped production bundle. Weekly automation is an in-process node-cron job (Mon 09:00 Central), idempotent (skips weeks already generated); on-demand generation and veracity recompute run via internal-token-gated HTTP routes. Dependencies: Anthropic via the Replit AI Integrations proxy (base URL + key in env), secrets AGENT_INTERNAL_TOKEN and SESSION_SECRET, and a Postgres database separate from development. Rollback is available via Replit checkpoints. Open readiness item: reliable weekly cron firing requires an always-on (Reserved VM) deployment, since autoscale scales to zero; on-demand works on any target. | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Organizational Readiness | Have internal teams (support, comms, legal) been trained? Is documentation complete? | Single-site internal tool; the ICQA Manager is both operator and primary reviewer, so there is no external support/legal/comms surface. Documentation comprises the system-prompts doc, the feature log, the eval harness, and this PRD; the review/approve flow has an operator guide. | ||||
| Launch & Rollout Strategy | Launch Approach | What is your launch approach? Pilot, AB test, or all users—who gets access and when? | Single-DC pilot (site T560), advisory-first: the agent's callouts merge into the existing deck while the ICQA Manager reviews veracity flags and approves each week, building trust before any broadening. No external users. | |||
| Scale Readiness | How will you ensure readiness for scale? How will you monitor initial volume and scale up? | Per-week processing with a 30-second cached deck payload, an idempotent cron, a bounded tool loop (≤8 iterations), and token caps keep per-run cost and latency predictable. Scaling to more sites means more weekly runs on the same pipeline; the eval harness and advisory veracity scale by adding fixtures/weeks. | ||||
| Go-to-Market Plan | Marketing / Training Assets | What assets (FAQ, demo, guides) will you prepare for external communication/marketing? | Internal only: the WBR deck itself (the artifact leadership reads), the presenter speaker notes, the system-prompt and feature-log documentation, and a short operator guide for the review/approve flow. | |||
| Stakeholder / Internal Comms | How will you communicate launch plans, progress, and outcomes internally? | Outputs land in the existing weekly leadership cadence (DBR/WBR). Changes to the agent are communicated through the same site reporting channel, and reviewer feedback is the internal improvement loop. | ||||
| Confirm Legal, Privacy & Risk Protocols | Data & Privacy | How do you handle and protect user data, including storage, privacy, and compliance? | Internal operational data only (demand, cancels, reason codes, workstation/zone, freight, team-member rosters). The only personal data is team-member names/logins used for contributor attribution; terminated team members surface as Z-logins and are never speculated on, and the SVIAITM automation account is excluded. Data is stored in Replit-managed Postgres with the production database separate from development; there is no external data sharing, and LLM calls go through the Replit AI Integrations proxy. | |||
| Policy & Compliance | Are content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain? | Content is governed by the narrative guard (no prescriptive/coaching language, no bare codes, no self-corrections) and the advisory veracity layer (grounding + judge) before approval, with an audit trail via agent_runs / agent_outputs (draft/edited/approved status plus persisted veracity). As an internal tool there is no external regulatory surface; the observational, non-prescriptive posture is itself a compliance guardrail. | ||||
| Define Success Metrics | Success Metrics | User/Business Metrics | What user metrics will indicate success? What business metrics will demonstrate value? | Reduction in “edit-required” weeks (how often the ICQA Manager must hand-rework the deck's story), reviewer approval rate and edit distance on agent callouts, time saved on WBR preparation, and consistency of the weekly leadership narrative. Downstream value: earlier visibility of multi-week drift (paths quietly above tolerance before they become the worst gap). | ||
| AI Metrics | How will you measure AI performance and accuracy? | Grounding pass rate and flagged-claim rate per run (broken out by finding kind), judge dimension scores (voice/accuracy/guard) and judge pass rate, fixture-suite pass/fail in CI, narrative-guard retry/drop rate, and tool-call counts per run. | ||||
| Monitor, Iterate & Improve | User Support & Feedback Plan | Support Channels | Where can users get support? Is escalation and ownership clear? | Internal: the ICQA Manager (operator) and the product/dev owner, with escalation through the site's normal tooling support path. Issues surface via the agent inbox / review feedback. | ||
| Feedback Workflow | How do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated? | The ReviewPanel captures like/dislike + reason-tag + note feedback per callout/note and at deck level. This is distilled server-side into a bounded lessons block that steers the next draft (Tone/Brevity/Phrasing → composer; Veracity/Accuracy/Relevance → investigator). Implicit feedback (callouts the reviewer dropped from an approved deck) also feeds lessons; a clear-checkpoint silences both sources on demand. Feedback never alters an already-served deck. | ||||
| Monitoring & Continuous Improvement | Monitoring Approach | What monitoring/logging is in place to spot operational/AI issues post-launch? | Structured, request-scoped logging (no console.log), agent_runs status and tool-call counts, advisory veracity persisted per draft, and the offline fixture gate in CI. Production runtime issues are visible via deployment logs; the cron logs each tick (skip / started / finished / failed) and never crashes the server. | |||
| Ongoing Improvement | How will you collect learnings, review performance, and update your system continuously post-launch? | The reviewer learning loop (lessons from explicit feedback + approved-diffs) continuously tunes the next draft; new failure modes become new golden fixtures (regression lock) and/or guard rules; live judge spot-checks catch voice/accuracy drift; and the clear-checkpoint lets the operator reset guidance when the data regime changes. | ||||




