Deliverable performance
Rubric-weighted correctness, requirements, tests, quality, and role-specific craft.
We measure how you ship with AI agents: planning, direction, verification, recovery, manual judgment, and final deliverable quality. Every competency is evidence-linked, so each score points back to something you actually did in the session.
Every candidate gets the same built-in AI tooling, interface, and time for the duration of the assessment, at no personal cost. The assessment fee covers all AI usage, so candidates who can't afford premium AI subscriptions aren't penalized. Where an assessment offers a choice of frontier models, every candidate sees the same list; recruiters can lock it to one model when a fair head-to-head on a single assistant matters more than choice.
One report, five separate layers
The report keeps performance, workflow evidence, integrity context, and human judgment visibly separate. Assessment creators choose a role-appropriate rubric before candidates begin; published cohorts keep that scoring contract fixed so results remain comparable within the assessment.
Rubric-weighted correctness, requirements, tests, quality, and role-specific craft.
Planning, prompting, verification, and deliberate use of the available workflow.
Whether enough attributable session evidence exists to support each interpretation. This is not a candidate-quality score.
Policy and session signals that may need human review. A flag is context, not an automatic rejection.
The hiring team's explicit advance, hold, or decline decision. AlgoArena never disguises this as an automated score.
Current weights and cut scores are expert-designed and awaiting calibration on real hiring cohorts. AlgoArena does not describe them as independently validated or universally predictive.
Each competency captures a different part of the AI-native engineering loop. Together, they create a readable picture of how you work with agents.
Whether what you shipped does what the task asked: correctness, core requirements, tests, and controlled scope.
What "good" looks like
A strong candidate ships the requested task first, then adds polish only when it supports the deliverable. Extra scope never compensates for broken core requirements.
Whether you read the task, break it into steps before directing the agent, and re-plan after a failure.
What "good" looks like
A strong candidate spends real time on the problem, notes an approach and the edge cases, and edits AI-generated plans instead of accepting them unread.
Whether you give the agent enough context, constraints, and acceptance criteria, refine the ask when output is weak, and match model spend to the task.
What "good" looks like
A strong candidate writes specific, context-rich prompts, reviews agent output, asks for focused changes, and rejects weak suggestions. They also match the model to the task: a premium model to think through a plan or a hard problem, a cheap, fast model to grind through the implementation it produced.
Whether you check the work: running tests, exercising the app, reading diffs, and challenging confident but wrong output.
What "good" looks like
A strong candidate verifies in their own style. Some read every diff before keeping it, others hammer the live preview and browser checks after each change. Both register. What reads poorly is neither, accepting output without ever reading or testing it.
Whether you pick the right tool for each step: plan, ask, code, search, terminal, browser checks, manual edits.
What "good" looks like
A strong candidate switches modes deliberately, delegates work that suits an agent, keeps ownership of the decisions, and spends tokens where they improve the deliverable.
Your assessment session is recorded and replayable. Scoring reads that record, and so does the company reviewing you.
The session is recorded and replayable
Editor activity, file changes, and timing are captured so a reviewer can replay how the work happened instead of guessing from the final code.
AI conversations
Every prompt you send, the reply it produced, and which suggested edits you kept or rejected.
Verification runs
Test runs, automated checks, and browser validation, along with their results.
Keystroke activity
Counts and timing only. We do not record which keys you press.
Integrity signals
Pastes, tab switches, and window focus changes.
Who can see it
The company that invited you to the assessment and authorized AlgoArena operators when needed to run or support the service. We do not sell assessment data; retention and other uses are described in the candidate notice and Privacy Policy.
Every competency is scored from evidence tied to the rubric, not surface-level metrics. Inflating keystrokes or running empty test commands will not improve your score. The best strategy is to direct AI well, verify the work, and ship the requested deliverable.
Every candidate gets built-in frontier AI at no personal cost, with the same interface, timer, and tooling for everyone. Some assessments let you pick from a short list of models; otherwise the recruiter locks one in. Do your AI work in that built-in chat, not an outside tool: we can only score prompting, reasoning, and verification we can see. Code pasted from an external assistant looks like unaided work and scores lower on those competencies.
Yes. We track code changes, keystroke activity (counts and timing, not which keys you press), AI prompts, test runs, tab switches, paste events, and timing. The session is recorded and replayable, and that record is what the scoring draws on. It is shared with the company that invited you. We do not sell assessment data; the candidate notice and Privacy Policy explain retention, service operations, and other permitted uses.
AlgoArena publishes the competencies on this page, and recruiters see the rubric-weighted breakdown with evidence references. The assessment creator can vary the rubric and weights for the role, so the report separates deliverable performance, AI-workflow competencies, evidence completeness, integrity status, and the recruiter decision. These are evidence-linked hiring aids, not a claim of psychometric validation.
Only when the assessment includes a baseline algorithmic task. The flagship signal is AI-native delivery: planning, directing agents, verifying output, recovering from mistakes, and shipping quality software.
Candidates get their assessment link by email when a company invites them. Hiring? See what this scoring looks like on your side.