Validra performance under load

How long each feature takes, and how fast it runs, as more people use it at the same time — measured on the live system.

Measured 30 September 2026

The one thing to know

The server has a single graphics card, and it produces about the same number of tokens per second in total whether one person is waiting or twenty-five.

So waits do not hold steady as people join — they stretch. Ten people share the speed one person would have had alone, and each waits roughly ten times longer.

A token is roughly three quarters of a word, so 10 tokens/sec is about 7 or 8 words a second appearing on screen.

Sequential — each person waits their turn before starting.
Parallel — everyone starts together and shares the machine.
p50 (usual) — the middle result; half were faster, half slower.
p90 / p95 — 9 in 10, and 19 in 20, waited no longer than this.
Slowest — the worst single wait recorded.
Avg tok/s — tokens per second a single request generated at.

Validra Assistant

Answering questions about an attached document

PeopleHowRunsPassed p50 (usual)p90p95Slowest Avg tok/s
5Sequential15/535 sec7m 35s8m 28s9m 21s27
Parallel15/535 sec3m 34s3m 37s3m 40s22
10Sequential110/1038 sec10m 31s10m 34s10m 37s40
Parallel220/2044 sec4m 14s9m 2s13m 50s26
25Sequential125/2550 sec1m 30s6m 7s19m 56s23
Parallel250/504m 28s12m 22s12m 30s12m 43s13

Its four test questions range from a one-line lookup to building a large table, so the usual and slowest figures sit far apart. Both are real.

AI Editor

Rewriting and generating text inside a document

PeopleHowRunsPassed p50 (usual)p90p95Slowest Avg tok/s
5Sequential15/534 sec34 sec34 sec34 sec29
Parallel15/537 sec39 sec39 sec39 sec26
10Sequential110/1034 sec36 sec39 sec42 sec33
Parallel110/1048 sec54 sec57 sec61 sec24
25Sequentialnot tested
Parallelnot tested

Document Chunking

Splitting an uploaded document into sections

PeopleHowRunsPassed p50 (usual)p90p95Slowest Avg tok/s
5Sequential210/1014 sec25 sec29 sec32 sec104
Parallelnot tested
10Sequential110/1014 sec15 sec18 sec20 sec104
Parallelnot tested
25Sequential125/2514 sec16 sec16 sec28 sec104
Parallel125/252m 28s3m 58s4m 9s4m 22s96

This one processes an uploaded document rather than answering a question, so its work is shaped differently from the others.

Data Analysis

Classifying an uploaded data file

PeopleHowRunsPassed p50 (usual)p90p95Slowest Avg tok/s
5Sequential210/1016 sec25 sec34 sec43 sec25
Parallel15/538 sec53 sec57 sec60 sec21
10Sequential110/1016 sec18 sec18 sec18 sec24
Parallelnot tested
25Sequential125/2516 sec17 sec18 sec20 sec22
Parallelnot tested

Report Master

Writing a full validation report

PeopleHowRunsPassed p50 (usual)p90p95Slowest Avg tok/s
5Sequential210/105m 5s5m 7s5m 7s5m 7s31
Parallel15/55m 13s5m 18s5m 18s5m 19s27
10Sequential110/1010m 6s10m 8s10m 9s10m 9s35
Parallelnot tested
25Sequentialnot tested
Parallelnot tested

Mixed load — every feature at once

People spread across all five features instead of queueing on one. This is closer to a normal working day than any single-feature run.

DatePeopleHowPassedWall clockAvg tok/sWho was doing what
2026-09-3025Parallel25/259m 34s71Validra Assistant 7, AI Editor 4, Document Chunking 6, Data Analysis 5, Report Master 3
2026-09-3025Sequential25/2511m 15s13Validra Assistant 8, AI Editor 3, Document Chunking 7, Data Analysis 4, Report Master 3
2026-09-3010Sequential10/104m 46s33Validra Assistant 1, AI Editor 1, Document Chunking 2, Data Analysis 3, Report Master 3
2026-09-2410Parallel10/105m 17s65Validra Assistant 3, AI Editor 2, Document Chunking 1, Data Analysis 2, Report Master 2

Percentiles are computed over every request in a cell. Where a cell holds only five requests, p90 and p95 sit close to the maximum and carry little more information than it does — they earn their keep at 25 people, not at 5.