Getting Started
- What is TestInvite?
- Build Your First Test
- Run Your First Assessment
- Taking the Assessment
- Viewing the Results
- Question Bank Overview
- Common Question Features
- Question Types
- Creating Questions
- Organizing Questions
- Content Blocks
- Roles & Access
- Media Library
- Import & Export
- Question Submissions
- Question Bank Schema
- Browsing Questions
- Cloning Questions
- Bulk Updating Questions
- Tests Overview
- My Tests
- Creating a Test
- The Test Editor
- Test Settings
- Sections & Pages
- Adding Questions
- Page Builders
- Test Profile
- Reporting
- Test Papers
- Analytics
- Publishing a Test
- Test Library
- Marketplace
- Tasks Overview
- Creating a Task
- Task Dashboard
- Steps
- Task Settings
- Candidates
- Test Sessions
- Sent Mails
- Proctoring
- Analytics
Tests
Test Analytics
Measure how your test performs as an instrument: score distributions and reliability at test level, difficulty and discrimination indices per question, and how to act on them to improve your test.
Updated 2026/07/13
Test analytics turn accumulated results into an assessment-quality report: how scores distribute, how consistent the test is as a measuring instrument, and — question by question — which items work, which are too easy or too hard, and which fail to separate strong candidates from weak ones.
Test-Level Analytics
- Score distribution — the spread of results: mean, median, standard deviation, and the shape of the curve. A healthy discriminating test shows spread; a distribution piled at 95% means the test is too easy for the audience (or leaked).
- Completion statistics — how many candidates started, finished, and how long they took, revealing timing problems (many candidates hitting the limit → the test is too long).
- Dimension breakdowns — average performance per scoring dimension across all candidates, showing which competencies your population is strong or weak in.
- Reliability — an internal-consistency estimate of the test as a whole. Low reliability means item scores don't hang together — usually a sign of mixed constructs or noisy items. Reliability figures need a reasonable sample; treat them as indicative until enough sessions accumulate.
Question Analytics
Each question gets two item-statistics that together tell you whether it earns its place:
- Difficulty index — the share of candidates who answered correctly. Near 1.0 = everyone gets it (too easy to inform ranking); near 0.0 = almost nobody does (too hard, or the question/key is broken). Most useful items sit in the middle band.
- Discrimination index — whether candidates who score well overall also do better on this question. Strongly positive = the item pulls in the same direction as the test. Near zero = the item adds noise. Negative = a red flag: strong candidates get it wrong more often than weak ones — almost always a miskeyed answer or an ambiguous prompt.
Acting on the Numbers
- Check every question with a negative discrimination first — verify the correct answer is keyed correctly.
- Review extreme-difficulty items: rewrite, replace, or keep deliberately (a warm-up item may be intentionally easy).
- Watch the score distribution against the test's purpose: a certification test can tolerate a high pass pile; a ranking test cannot.
- Re-run analytics after each significant cohort — item statistics stabilize as sessions accumulate.
Item statistics are population-relative: the same question is “hard” for juniors and “easy” for seniors. Interpret analytics against the audience that actually took the test, and compare cohorts rather than absolute values when the audience changes.