Question Bank Scoring

Evaluation & Grading

Not every answer can be checked automatically right or wrong — Manual Evaluation, Rubric Scoring, and AI Evaluation are the three ways a question's success rate gets produced by judgment instead of a formula.

Published 2026/07/30

Multiple Choice, Numeric, and similar question types score themselves automatically — the answer is either right, wrong, or partially credited by a fixed rule. Open-ended types (Long Answer, Code, Audio, Video, and similar) have no such rule: someone or something has to judge the response and produce a success rate. TestInvite offers three ways to do that, and they can coexist on the same question.

The Three Evaluation Methods

  • Manual Evaluation — a reviewer enters a single score directly, with no structure imposed. The simplest method, and the fallback when neither a rubric nor AI evaluation is configured.
  • Rubric Scoring — a reviewer works through a structured set of criteria and levels instead of freehand, so multiple reviewers score the same response consistently.
  • AI Evaluation — AI computes the score automatically from a grading prompt you write, for the question types that support it — no reviewer needed unless you choose to check its work.

A fourth method, per-dimension scoring, is really a different axis rather than a competing evaluation method — it's about scoring several competencies on one response independently rather than producing one overall score. See Scored Dimensions for that.

What Happens When More Than One Is Available

When a question has AI Evaluation enabled, the AI evaluates first automatically. If a reviewer then enters a score — through the per-dimension rows, the rubric, or a plain manual entry, whichever method is active on the question — that manual entry takes precedence and suppresses the AI's result for that response. In short: AI fills the gap until a human weighs in, and a human's judgment always wins.

See Manual Evaluation, Rubric Scoring, and AI Evaluation below for how each one works.

Manual Evaluation is a single score a reviewer enters directly, with no rubric or AI involved — the default way to grade a question when nothing more structured is set up.

How rubric scoring works in TestInvite — what rubrics are, which question types support them, and an overview of the four rubric types.

Score open-ended answers automatically with a large language model: how AI evaluation works per question type, how to write an effective grading prompt, and how reviewers confirm or adjust the results.