Skip to main content

How AI-Native Scoring Works

We measure how you ship with AI agents: planning, direction, verification, recovery, manual judgment, and final deliverable quality. Every competency is evidence-linked, so each score points back to something you actually did in the session.

Level Playing Field

Every candidate gets the same built-in AI tooling, interface, and time for the duration of the assessment, at no personal cost. The assessment fee covers all AI usage, so candidates who can't afford premium AI subscriptions aren't penalized. Where an assessment offers a choice of frontier models, every candidate sees the same list; recruiters can lock it to one model when a fair head-to-head on a single assistant matters more than choice.

One report, five separate layers

A score is not a hiring decision.

The report keeps performance, workflow evidence, integrity context, and human judgment visibly separate. Assessment creators choose a role-appropriate rubric before candidates begin; published cohorts keep that scoring contract fixed so results remain comparable within the assessment.

01

Deliverable performance

Rubric-weighted correctness, requirements, tests, quality, and role-specific craft.

02

AI-workflow competencies

Planning, prompting, verification, and deliberate use of the available workflow.

03

Evidence completeness

Whether enough attributable session evidence exists to support each interpretation. This is not a candidate-quality score.

04

Integrity status

Policy and session signals that may need human review. A flag is context, not an automatic rejection.

05

Recruiter decision

The hiring team's explicit advance, hold, or decline decision. AlgoArena never disguises this as an automated score.

Current weights and cut scores are expert-designed and awaiting calibration on real hiring cohorts. AlgoArena does not describe them as independently validated or universally predictive.

Five Competencies We Evaluate

Each competency captures a different part of the AI-native engineering loop. Together, they create a readable picture of how you work with agents.

Problem Solving & Deliverable Quality

Whether what you shipped does what the task asked: correctness, core requirements, tests, and controlled scope.

What "good" looks like

A strong candidate ships the requested task first, then adds polish only when it supports the deliverable. Extra scope never compensates for broken core requirements.

Evidence used:The graded deliverableTest runs and their resultsAutomated checks against your buildWorkspace snapshots taken as you go

Planning & Decomposition

Whether you read the task, break it into steps before directing the agent, and re-plan after a failure.

What "good" looks like

A strong candidate spends real time on the problem, notes an approach and the edge cases, and edits AI-generated plans instead of accepting them unread.

Evidence used:Your prompts and the replies they producedSession replay of how you workedThe written session summary

AI Direction & Communication

Whether you give the agent enough context, constraints, and acceptance criteria, refine the ask when output is weak, and match model spend to the task.

What "good" looks like

A strong candidate writes specific, context-rich prompts, reviews agent output, asks for focused changes, and rejects weak suggestions. They also match the model to the task: a premium model to think through a plan or a hard problem, a cheap, fast model to grind through the implementation it produced.

Evidence used:Your prompts and the replies they producedThe written session summary

Verification & Iteration

Whether you check the work: running tests, exercising the app, reading diffs, and challenging confident but wrong output.

What "good" looks like

A strong candidate verifies in their own style. Some read every diff before keeping it, others hammer the live preview and browser checks after each change. Both register. What reads poorly is neither, accepting output without ever reading or testing it.

Evidence used:Test runs and their resultsAutomated checks against your buildSession replay of how you workedYour prompts and the replies they produced

Agentic Workflow Autonomy

Whether you pick the right tool for each step: plan, ask, code, search, terminal, browser checks, manual edits.

What "good" looks like

A strong candidate switches modes deliberately, delegates work that suits an agent, keeps ownership of the decisions, and spends tokens where they improve the deliverable.

Evidence used:Session replay of how you workedYour prompts and the replies they producedWorkspace snapshots taken as you go

What Is Recorded

Your assessment session is recorded and replayable. Scoring reads that record, and so does the company reviewing you.

The session is recorded and replayable

Editor activity, file changes, and timing are captured so a reviewer can replay how the work happened instead of guessing from the final code.

AI conversations

Every prompt you send, the reply it produced, and which suggested edits you kept or rejected.

Verification runs

Test runs, automated checks, and browser validation, along with their results.

Keystroke activity

Counts and timing only. We do not record which keys you press.

Integrity signals

Pastes, tab switches, and window focus changes.

Who can see it

The company that invited you to the assessment and authorized AlgoArena operators when needed to run or support the service. We do not sell assessment data; retention and other uses are described in the candidate notice and Privacy Policy.

What We Don't Measure

Algorithm memorization as the main signal
LeetCode pattern matching
Typing speed or raw keystroke count
Prompt volume without useful progress
How fast you finish, engagement matters more
Trick questions or gotcha puzzles

Frequently Asked Questions

Can I game the score?

Every competency is scored from evidence tied to the rubric, not surface-level metrics. Inflating keystrokes or running empty test commands will not improve your score. The best strategy is to direct AI well, verify the work, and ship the requested deliverable.

What AI tools can I use?

Every candidate gets built-in frontier AI at no personal cost, with the same interface, timer, and tooling for everyone. Some assessments let you pick from a short list of models; otherwise the recruiter locks one in. Do your AI work in that built-in chat, not an outside tool: we can only score prompting, reasoning, and verification we can see. Code pasted from an external assistant looks like unaided work and scores lower on those competencies.

Is my code tracked? What data do you collect?

Yes. We track code changes, keystroke activity (counts and timing, not which keys you press), AI prompts, test runs, tab switches, paste events, and timing. The session is recorded and replayable, and that record is what the scoring draws on. It is shared with the company that invited you. We do not sell assessment data; the candidate notice and Privacy Policy explain retention, service operations, and other permitted uses.

How can I understand my score?

AlgoArena publishes the competencies on this page, and recruiters see the rubric-weighted breakdown with evidence references. The assessment creator can vary the rubric and weights for the role, so the report separates deliverable performance, AI-workflow competencies, evidence completeness, integrity status, and the recruiter decision. These are evidence-linked hiring aids, not a claim of psychometric validation.

Do you measure algorithm knowledge?

Only when the assessment includes a baseline algorithmic task. The flagship signal is AI-native delivery: planning, directing agents, verifying output, recovering from mistakes, and shipping quality software.

Candidates get their assessment link by email when a company invites them. Hiring? See what this scoring looks like on your side.