← All capstone projects

Education

Academic Certificate Helper

Built by JoeyTri Cohort 9 Education / test preparation

Academic Certificate Helper is an AI-driven prep product aimed at Southeast Asian students pursuing Western education credentials, starting with English proficiency exams. It generates mock tests, scores performance, highlights weak areas, and exposes results for both learners and tutors to review. The pitch positions the product as a cheaper, more trustworthy alternative to ad hoc AI practice and expensive prep cycles.

The problem

Southeast Asian students pursuing Western education credentials face a trust problem at the center of every study session. Distrust of the AI grade is the most severe and frequent pain: frontier LLMs match human examiner bands only 35–65% of the time and are 15–20% biased against ESL writers — who are the entire user base. Compounding it are the expensive retry loop (heavy spend on classes, mocks, and $180+ fees, then a blind restart after a missed score) and mock-fidelity doubt (self-prompted AI mocks with no guarantee they match the real exam). Rigid, lengthy mocks don't fit daily life, and after a miss there's often no clear explanation of what to fix.

The solution

Academic Certificate Helper is an AI prep product for Southeast Asian students, starting with English proficiency exams. It generates exam-faithful mock tests, scores performance, highlights weak areas, and exposes results for both learners and tutors to review. The offering is phased: a free layer of bite-size, exam-faithful daily mini-mocks across all five certificates as the habit and lead-gen entry point; a paid depth layer of calibrated AI grading of speaking and writing plus longitudinal flaw detection — "know exactly what to fix" — that free tools don't provide; and a later tutor marketplace fed by that flaw history. It's positioned as a cheaper, more trustworthy alternative to ad hoc AI practice and expensive prep cycles.

How it works

The product's core AI opportunities, ranked by severity, are calibrated, ESL-fair grading (band plus rationale, calibration evidence, and a bias monitor — the existential bet), longitudinal flaw detection that mines graded history into a prioritized "what to fix," and accuracy-assured mock generation using blueprint-grounded generation, validators, and an independent cross-provider LLM-judge for exam fidelity. Additional layers include on-demand bite-size mock generation and personalized targeted practice aimed at detected weaknesses. Some elements are deliberately kept non-AI — streak mechanics, free-versus-paid positioning, minor-consent flows, and marketplace matching and payments. The current prototype is IELTS-only, with no backend, auth, or persistence, running on a thin proxy.

Who it's for

The product is primarily B2C, for external learners: a Southeast Asian academic preparing for one of five Western certificates (SAT, IELTS, TOEFL, PTE, GMAT), ESL by definition, mobile-first, and already studying with AI daily. Three variants shift the journey — SAT learners are under-18 minors with a parent co-payer and a legally required consent step; IELTS/TOEFL/PTE learners are the high-volume wedge (Vietnam-first, ages 16–22); GMAT learners are the highest-spend but lowest-priority adults. It becomes B2B2C in the marketplace phase, when tutors join as a supply side and consume AI-derived student previews rather than the AI features directly. Launch focuses on the SEA-6 region, Vietnam-first.

Why it matters

AI study behavior is already normalized — 86% of students use AI, 54% weekly — and exams have gone digital, short, and adaptive, so real tests now resemble the mini-mocks. AI-first testing is institutionally accepted, with 1 in 5 US international applicants using the Duolingo English Test in 2025. The broad test-prep market grows at roughly 13.5% CAGR, with an estimated TAM around $2B/yr for these five exams. The defensible advantage is a three-pillar bundle no competitor combines: exam-faithful mini-mocks across all five certs, longitudinal flaw-detection history, and a tutor marketplace fed by that history. Each pillar alone is replicable; the moat is the combination plus the longitudinal flaw data. The business is pre-seed, pre-launch, phased free habit → grading and flaw detection → marketplace and monetization.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Joey (Tri) Nguyen
Your Product:Languages Proficiency Test helper
Your Industry:Education
Date:May 4, 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
Our business operates in AI-assisted exam preparation — the EdTech discipline at the intersection of generative AI and high-stakes standardized testing — and creates value by encoding exam fidelity and each learner's longitudinal performance into a structured, machine-readable profile that AI generation, AI grading, and human tutors can all read and act on. To deliver this value, our product offers free, bite-size, exam-faithful mini-mocks across all five Western certificates (SAT, IELTS, TOEFL, PTE, GMAT), integrated with calibrated AI grading and longitudinal flaw detection that read from each learner's history, meeting the needs of candidates preparing for high-stakes admissions and migration exams across receptive skills, productive speaking and writing, and the daily habit that has to fit around their lives. These customers are primarily Southeast Asian study-abroad, immigration, and admissions candidates — ESL by definition — preparing with fragile or absent feedback on what is actually holding their score back, who experience the expensive retry loop of repeated ~$180+ exam fees and re-prep, rigid full-length mocks that don't fit daily life, and self-prompted AI practice with no fidelity guarantee (frontier LLMs match human examiner grade bands only 35–65% of the time) as they try to reach their target score without burning successive attempts and tutoring spend blindly guessing at what to fix. To address these pain points, we propose an AI-powered solution — the Accuracy-Assured Mock & Flaw-Detection Engine — that generates exam-faithful mocks, grades every skill against a calibrated per-exam standard, and diagnoses each learner's recurring weaknesses in real time, surfacing them with explanations and a prioritized "what to fix," delivering value to the learner (less wasted money and prep time, a clear path to the target score), the payer (the learner or, for under-18 SAT candidates, the parent — visible return on prep spend through fewer retakes), and the wider ecosystem of tutors and admitting institutions who receive pre-diagnosed, better-prepared candidates.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Industry: Education — specifically EdTech / standardized-test preparation, narrowed to AI-assisted exam prep for the 5 Western certificates (SAT, IELTS, TOEFL, PTE, GMAT), launching in Southeast Asia (SEA-6: Vietnam, Indonesia, Thailand, Philippines, Malaysia, Singapore), Vietnam-first.Please leave this area blank. This space is for the Instructor to provide you with feedback.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?Headwinds: Unreliable AI grading: frontier LLMs matched human grade bands only 35–65% of the time (Cambridge, May 2026) — existential for an "accuracy-assured" product. ESL grading bias: 15–20% score discrepancy against non-native writers — your entire user base is ESL. Minor-data compliance: Vietnam PDPL (in force 2026-01-01) requires dual consent for SAT's under-18 takers — launch-blocking for that segment. Test owners ship the same feature: British Council "IELTS Ready" and ETS "TOEFL TestReady" already give AI mock feedback. The floor is free: Khan Academy owns official free SAT prep; free ChatGPT mocks circulate. Can't charge for generic practice. GMAT structurally shrinking: −19% YoY (93,196 exams, TY2025). Tailwinds (opportunities): AI study behavior is normalized: 86% of students use AI; 54% weekly — near-zero adoption barrier. Exams went digital/short/adaptive — real tests now look like your mini-mocks (high practice-to-test fidelity). AI-first testing is institutionally accepted: 1 in 5 US international applicants used Duolingo English Test scores in 2025; accepted by 6,000+ programs. Rising test fees widen the prep-value gap; SEA outbound mobility growing, VN-led; VC validation in-geography: ELSA ($23M Series C), Prep ($7M Series A).
What is the projected growth rate of your target market segment over the next 3-5 years?Projected growth (3–5 yr): Broad industry proxy: ~13.5% CAGR (global test-prep market $11.61B in 2025 → 2033; Verified Market Reports, SECONDARY). TAM (your 5 exams): ~$2B/yr (range $1.2–2.9B, ESTIMATE). Honest caveat: no published source gives a clean CAGR for this exact segment (5 Western exams × SEA-6). The picture is divergent: IELTS/language exams growing with SEA mobility (e.g. Indonesia +29% students abroad), GMAT declining −19%. The growth thesis rests on SEA demand expansion, not the shrinking exams.
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Pre-seed / early startup. The docs describe a prototype (IELTS-only, no backend/auth/persistence, runs on a thin proxy) plus a planning engagement — this is plans-and-specs, pre-revenue, pre-launch. Phasing is B1 (free habit) → B2 (grading + flaw detection) → B3 (marketplace + monetization).
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)Free (always): bite-size, exam-faithful daily mini-mocks across all 5 certs — the habit/lead-gen layer. Paid depth (B2): calibrated AI grading of speaking/writing + longitudinal flaw detection ("know exactly what to fix") — the thing free tools don't do. Marketplace (B3): tutor discovery where tutors see a student's pre-diagnosed strengths/weaknesses (the C7 preview moat).
Who is your primary customer base (B2B, B2C, B2B2C)?B2C primarily (learners — students, study-abroad/immigration candidates; SAT segment skews minors), becoming B2B2C in B3 when the two-sided tutor marketplace opens (learners + tutor supply side).
DifferentiatorsWhat are the key differentiators for your company?Core unfair advantage (BD-04 §8, RE-04 obs. 1): the three-pillar bundle is unoccupied — no competitor combines all three: Exam-faithful AI mini-mocks across all 5 certs (blueprint-grounded + validator + cross-provider LLM-judge accuracy pipeline — your D7/D9 architecture is the credibility engine here). Longitudinal flaw-detection history — calibrated, per-user "what to fix" over time. Incumbents give one-shot feedback; none track flaws longitudinally. Tutor marketplace fed by that flaw history (C7 preview) — tutors see diagnosed weaknesses before the first lesson; Preply/italki have undifferentiated supply. Why it's defensible: singly, each pillar is replicable (free tools, marketplaces, test-owner feedback all exist). The moat is the combination + the longitudinal flaw data feeding tutor previews — the asset incumbents can't quickly copy. Secondary differentiators: Calibrated, ESL-aware grading — turning the H1/H2 weaknesses (unreliable, ESL-biased AI grading) into a trust advantage via a real calibration mechanism. Compliance-as-trust — PDPL-compliant minor-data handling for the SAT segment (turns a constraint into a signal). Habit over cramming — bite-size daily mocks fitting the digital/short/adaptive shape real exams now take.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?[Insert your response here]
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?[Insert your response here]
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?[Insert your response here]
User Value MapTarget PersonaWho is your AI product / feature for?External end-users — the learners. Primary persona: a Southeast Asian academic preparing for one of the five Western certificates (SAT, IELTS, TOEFL, PTE, GMAT), ESL by definition, mobile-first, and already studying with AI daily. Three variants shift the journey's edges — SAT learners are under-18 minors with a parent as co-payer and a legally required consent step; IELTS/TOEFL/PTE learners are the high-volume wedge (Vietnam-first, ages 16–22); GMAT learners are adults with the highest spend but lowest priority. A secondary, non-AI-facing user, the tutor, appears in the marketplace phase and consumes AI-derived student previews rather than the AI features directly.
Journey Map (current-state)What is the typical (happy-path) journey for your target persona?(1) Discover a free, short, exam-faithful daily mock and install it. (2) Onboard — create an account and pick certificate(s); minors complete guardian consent. (3) Take a first bite-size receptive mock that visibly looks like the real exam. (4) Get an instant objective score and section breakdown, saved to history. (5) Build a daily habit through streaks and reminders — small reps instead of one long mock. (6) Grade the hard skills — record speaking or write essays and get a calibrated AI band with a rationale. (7) See a prioritized "what to fix" flaw report mined from their own history. (8) Practice against those weaknesses, re-test, and watch the flaw list shrink. (9) On a wall, reach for a human tutor who can preview their diagnosed gaps. (10) Sit the real exam with a built habit, graded skills, and a worked flaw list.
Pain-pointsWhere does the user experience friction or unmet needs, and which pain-points are most frequent and severe?Ranked by frequency × severity: (1) Distrust of the AI grade — bands feel wrong; frontier LLMs match human examiners only 35–65% of the time and are 15–20% biased against ESL writers, who are the entire user base [most severe]. (2) The expensive retry loop — heavy spend on classes/mocks/~$180+ fees, then a blind restart if the score is missed. (3) Mock-fidelity doubt — self-prompted AI mocks have no guarantee they match the real exam. (4) Rigid, lengthy mocks don't fit daily life. (5) Habit drop-off / broken streaks, worse on higher-effort productive tasks. (6) "Why pay?" when official prep (Khan Academy SAT, British Council IELTS Ready) is free. (7) No clear explanation of what to fix after a miss. (8) Consent friction for under-18 SAT learners. (9) No human help when AI alone isn't enough. Most frequent-and-severe are #1–#3, which recur at the trust-defining moment of every session.
AI OpportunitiesWhich pain-points can GenAI address, ranked most severe/frequent first?LLM-addressable, ranked: (1) Calibrated, ESL-fair AI grading — band + rationale + calibration evidence + bias monitor (solves #1; the existential bet). (2) Longitudinal flaw detection — LLM mines graded history into a prioritized, explained "what to fix" (solves #2 and #7; the core value moment). (3) Accuracy-assured mock generation — blueprint-grounded generation + validators + an independent cross-provider LLM-judge for exam fidelity (solves #3). (4) On-demand bite-size mock generation — unlimited short, daily, exam-shaped tasks at near-zero marginal cost (solves #4). (5) Personalized targeted practice — LLM generates the next items aimed at detected weaknesses (solves #2 and #7). (6) AI metadata/difficulty tagging to power adaptive practice (solves #4 and #7). (7) Disputed-grade explanation that bridges to human review (solves #1 and #9). Deliberately NOT GenAI: streak mechanics (#5), free-vs-paid positioning (#6), minor-consent flows (#8), and marketplace matching/payments (#9).
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.[Insert your response here]
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.[Insert your response here]
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?[Insert your response here]Please leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?[Insert your response here. Include link to visuals as appropriate.]
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?[Insert your response here. Include link to visuals / prototype as appropriate.]
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?[Insert your response here]
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?[Insert your response here]
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?[Insert your response here]
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?[Insert your response here]Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.[Insert your response here]
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?[Insert your response here]
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)[Insert your response here]
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?[Insert your response here]
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.[Insert your response here]
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?[Insert your response here]
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?[Insert your response here]
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.[Insert your response here]
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)[Insert your response here]
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?[Insert your response here]
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?[Insert your response here]
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?[Insert your response here]
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?[Insert your response here]
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?[Insert your response here]
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?[Insert your response here]
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?[Insert your response here]Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?[Insert your response here]
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?[Insert your response here]
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?[Insert your response here]
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?[Insert your response here]
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?[Insert your response here]
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?[Insert your response here]
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?[Insert your response here]
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?[Insert your response here]
AI MetricsHow will you measure AI performance and accuracy?[Insert your response here]
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?[Insert your response here]
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?[Insert your response here]
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?[Insert your response here]
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?[Insert your response here]
Download the .xlsx ↓