We measure how you ship with AI agents: planning, direction, verification, recovery, manual judgment, and final deliverable quality.
Every candidate gets the same built-in frontier-grade AI assistant for the duration of the assessment. The assessment fee covers all AI usage. Candidates who can't afford premium AI subscriptions aren't penalized. You compete on skill, not wallet.
Each dimension captures a different part of the AI-native engineering loop. Together, they create a readable picture of how you work with agents.
We look at how you break down a problem before directing AI. Do you read the requirements carefully, create a plan, and adjust it when the agent hits friction?
What "good" looks like
A strong candidate takes real time to read the problem, jot down approach notes, and identify edge cases before touching the IDE. They create or edit AI-generated plans rather than blindly proceeding.
We evaluate how clearly you direct AI agents: context, constraints, acceptance criteria, follow-up questions, and whether you refine the ask when output is weak.
What "good" looks like
A strong candidate writes specific, context-rich prompts. They review agent output, ask for focused changes, reject weak suggestions, and explain the intent behind the work.
We measure how actively you test, iterate, and verify agent output. Verification counts on both channels: reading the AI's diffs, or exercising the result — running the app, clicking through your UI, and using browser checks. Do you challenge confident but wrong AI output?
What "good" looks like
A strong candidate verifies in their own style: some read every diff before keeping it, others hammer the live preview and browser checks after each change. Both register. What reads poorly is neither — accepting output without ever reading or testing it.
We track whether you use plan, ask, code, search, commands, browser validation, and manual edits with judgment instead of letting the agent wander.
What "good" looks like
A strong candidate switches modes deliberately, delegates appropriate work to agents, keeps ownership of decisions, and spends tokens where they improve the deliverable.
We evaluate the final deliverable: core requirements, correctness, code structure, tests, browser behavior where relevant, scope control, and useful product craft.
What "good" looks like
A strong candidate ships the requested task first, then adds helpful polish only when it supports the deliverable. Extra scope never compensates for broken core requirements.
We track evidence tied to the assessment rubric, not surface-level metrics. Artificially inflating keystrokes or running empty test commands won't improve your score. The best strategy is to actually direct AI well, verify the work, and ship the requested deliverable.
During AlgoArena assessments, every candidate gets the same built-in frontier-grade AI assistant. You don't need your own subscriptions. Everyone gets the same capabilities, so it's a level playing field. Do your AI-assisted work in that built-in chat, not an outside tool — we can only score the prompting, reasoning, and verification we can actually see. Code pasted in from an external assistant looks identical to code you wrote yourself with no AI collaboration at all, and will score lower on those dimensions than the same help asked for here.
Yes, we track code changes, keystroke activity (counts and timing, not which keys you press), AI prompts, test runs, tab switches, paste events, and timing. This data is used for session replay and scoring. It's shared with the company that invited you to the assessment. We don't sell your data or use it for any other purpose.
CodeSignal's scoring criteria are largely opaque, so candidates and recruiters don't know what's being measured. AlgoArena publicly documents our evaluation dimensions (this page), and recruiters see the full weighted breakdown with raw signals. We believe transparency leads to better hiring decisions.
Only when the assessment includes a baseline algorithmic task. The flagship signal is AI-native delivery: planning, directing agents, verifying output, recovering from mistakes, and shipping quality software.
Candidates get their assessment link by email when a company invites them. Hiring? See what this scoring looks like on your side.