Elevorix logo Elevorix AI
Public Evaluation Standard

You know every question before you walk into the defense.

Six rubric criteria, published openly. If you can answer all six from your own system with real evidence, you are ready. No hidden standards, no surprises.

Why we publish this: Most programs hide their rubrics. We publish ours because if you have genuinely built what the diploma asks for, you should be able to pass on the first attempt. Use the self-assessment checkboxes in each criterion below to track what you can already answer with evidence from your own system.

Before You Read the Rubric

How the defense works

The capstone defense is a structured session — not a presentation, not an exam. You explain your system to an evaluator who asks questions from the rubric below. All six areas must be addressed.

Typical duration

30–50 minutes depending on track and system complexity. Questions are drawn from all six rubric areas.

Re-defense policy

Insufficient areas receive written feedback. You may revise and re-defend within an agreed timeline. The certificate is issued only after all criteria meet the standard.

The Six Rubric Criteria

What evaluators assess — and how

Each card below shows the evaluator questions, what a passing answer looks like, and what partial or insufficient responses look like.

Self-assessment — 0 / 24 questions you can answer with real evidence from your own system
Ready to defend
01 · Design Rationale

Why did you make this design decision?

0/4

Evaluator questions

  • Why did you choose this chunking or data partitioning strategy?
  • What alternatives did you consider before choosing this architecture?
  • Why did you select this model or tool for this specific component?
  • What trade-offs did your design choices create, and were they acceptable?

Score levels

Pass Names 2+ alternatives considered, explains the specific trade-off that led to the chosen approach, links the choice to a measurable system property (speed, quality, cost).
Partial Explains the chosen approach correctly but cannot articulate why alternatives were rejected. Design decisions feel intuitive, not reasoned.
Insufficient Cannot explain design decisions beyond "I followed a tutorial" or "it worked." No evidence of deliberate architectural thinking.
02 · Evaluation Discipline

How did you measure quality?

0/4

Evaluator questions

  • What metrics did you define before building, not after?
  • How did you evaluate retrieval quality in your RAG or NLP system?
  • How do you know your system is producing correct outputs, not just plausible ones?
  • What would cause your evaluation score to drop, and how would you detect it?
Track note: Automation track — workflow success rate, node error rate, execution time. NLP/RAG — precision, recall, MRR, RAGAS scores. CV — mAP, IoU, F1. The criterion is universal; the metrics are domain-specific.

Score levels

Pass Shows documented evaluation results (precision, recall, MRR, mAP, or equivalent) collected during development. Metrics were defined before deployment, not after reviewing outputs.
Partial Knows the right metrics and can describe what good results look like, but evaluation was informal. Numbers exist but collection method is unclear.
Insufficient Relies on qualitative feel ("it looks good") with no documented evaluation framework. Cannot name a specific metric for their system domain.
03 · Failure Handling

What happens when something breaks?

0/4

Evaluator questions

  • What happens when the tool or API your system depends on is unavailable?
  • How do you detect when your model is producing hallucinations or low-confidence outputs?
  • What is your fallback when retrieval returns no relevant results?
  • What are the known failure modes of your system, and how are they handled?

Score levels

Pass Identifies 3+ named failure modes with corresponding handling logic (fallback response, circuit breaker, confidence threshold, logging and alert). Failure handling is implemented, not theoretical.
Partial Identifies failure modes correctly but handling is incomplete — only some failure paths are addressed. Shows awareness of the problem without full implementation.
Insufficient Cannot name failure modes. System assumes inputs and dependencies are always available. No graceful degradation logic present.
04 · Deployment Evidence

Is this deployed, not just running locally?

0/4

Evaluator questions

  • Where is your system running? Show the endpoint or demo.
  • What environment does your system run in production?
  • How is configuration managed between development and deployment?
  • How would someone else run this system from your documentation?

Score levels

Pass Live endpoint, Docker container, n8n workflow instance, or deployed inference API with documented setup steps. Someone else can follow the README and run the system without asking questions.
Partial System works and is packaged for deployment but is not currently deployed to an accessible environment. Documentation is present but incomplete.
Insufficient System only runs on the learner's local machine with manual setup steps. No documented path for another person to reproduce the environment.
05 · Scalability Awareness

How does this behave under real load?

0/4

Evaluator questions

  • How would this system perform with ten times the current data volume?
  • Where is the performance bottleneck in your current architecture?
  • What would you change first if you needed to handle 1,000 concurrent users?
  • What is the cost profile of running this system at scale?

Score levels

Pass Identifies the current bottleneck by name (vector index size, LLM API rate limit, CPU at inference, workflow concurrency), explains a specific scaling strategy for that bottleneck, and estimates rough cost at 10× volume.
Partial Understands that scalability is a concern and can name general strategies (caching, batching, queue), but cannot pinpoint the specific bottleneck in their own system.
Insufficient Has not thought about performance beyond getting the system to work. Cannot describe any scenario where the current system would degrade under increased load.
06 · Communication Clarity

Can you explain this to a non-technical stakeholder?

0/4

Evaluator questions

  • Explain what your system does in two sentences without technical jargon.
  • What are the known limitations of this system that a business user should know?
  • If this system fails silently, how would an end user know something is wrong?
  • What would you change about this system if you were deploying it for a paying client?

Score levels

Pass Gives a clear two-sentence non-technical description. Lists at least 3 real limitations with business impact. Describes how end-user failure looks different from technical failure.
Partial Can explain what the system does but lapses into jargon when pressed. Limitations are acknowledged but framed as technical edge cases rather than user experience concerns.
Insufficient Cannot simplify their explanation beyond technical terms. Does not treat limitations as relevant information a user or client would need. Shows no stakeholder perspective.
Defense Red Flags

If you hear yourself saying any of these, stop.

These are the exact phrases that cause a defense to fail. They are drawn directly from the Insufficient descriptions above. Recognising them before the session is the point of this page.

Design Rationale
"I followed a tutorial" / "it worked so I kept it"

No evidence of deliberate architectural choice. You need to name what you considered and rejected — not just what you picked.

Evaluation
"It looks good" / "I tested it manually"

Qualitative feel is not evaluation. You need a named metric, a number, and proof the metric was measured before you finished building — not after.

Failure Handling
"I haven't thought about what happens if it breaks"

Every production system fails. If you cannot name three failure modes and their handlers, your system is not production-ready.

Deployment
"It only runs on my machine" / "you'd need to ask me to set it up"

A system that only one person can run is not deployed. Someone else must be able to reproduce your environment from your documentation alone.

Scalability
"I haven't thought about performance beyond getting it to work"

You must be able to name the specific bottleneck in your own architecture — not a general statement about caching or queues.

Communication
"I can't explain it without using technical terms"

If a business user cannot understand what your system does and when it fails, you have not understood it well enough yourself.

Found gaps in your self-assessment?

That is the point of this rubric — to surface what needs more work before the defense, not during it. Message a mentor directly on WhatsApp with the criteria you are unsure about and get specific guidance on what to build next.

WhatsApp Advisor