AI-Native Assessments Need Evidence, Not Theater
Why the new assessment scoring model focuses on competencies, browser evidence, verification, and artifacts a reviewer can see instead of pretending AI disappeared.
The old assessment question was simple: did the candidate pass the tests?
That is still important, but it is no longer enough. In 2026, candidates are working with AI tools, browser previews, generated tests, agents, and fast feedback loops. A useful assessment has to measure judgment inside that workflow instead of pretending the tools are not there.
The scoring model changed
The current assessment scoring profile reports five public competencies:
- Problem Solving & Deliverable Quality
- Planning & Decomposition
- Prompting & Communication
- Verification & Iteration
- Agentic Workflow Autonomy
Those dimensions make the report more explainable. A hiring team can see whether a candidate planned well, verified meaningfully, directed AI clearly, and delivered working software. A candidate can understand what was actually measured.
Five is a deliberate number. One overall score is easy to publish and hard to argue with, which is exactly the problem, because it hides the disagreement instead of showing it. Splitting the judgment into a small set of named axes forces the report to say where a candidate was strong and where they were not, which gives a reviewer somewhere to push back.
Evidence beats vibes
For UI and project-style work, final code is not the whole artifact. The assessment flow now has room for browser validation, DOM snapshots, viewport checks, console findings, prompt evidence, and reviewer notes. None of it helps unless a reviewer can see what was checked, because a green result with no context behind it is a badge taken on faith.
That is the difference between theater ("the candidate used AI") and evidence ("the candidate used AI, inspected the result, caught the broken mobile state, and fixed it").

Two candidates can hand in the same passing work, one having directed the tool and fixed the failure the preview caught, the other having accepted every suggestion unread. Only the trail tells them apart, and How to Assess Developers Who Code With AI works through a case like that step by step.
Why not just turn the tools off
The answer that looks cleanest is to ban the tools and grade the code the old way. It is tempting because it removes the hard measurement problem. It also measures a workflow the job no longer uses. A no-AI assessment tells a hiring team how a candidate performs in an environment they will never work in again, which is a strange thing to select on.
There are real cases for an AI-off baseline, and the presets support one, which is why we kept it. But it should be a deliberate choice for a specific reason, not the default hiding place for a product that found AI too messy to score. The Problem Was Never Vibecoding makes the longer case against banning the tools.
Presets without changing the language
Different roles need different weights. A front-end build can care more about browser evidence. A classic no-AI baseline can set prompting to zero. A debugging task can care more about iteration and verification.
But the public report should not invent a new language for every assessment. The five competencies stay stable while role presets adjust weights underneath.
That stability earns its keep with the people reading dozens of reports a week. A reviewer learns one report shape, and every assessment after that is faster to judge because the axes never move even when their weights do.
When evidence becomes noise
Buried in all of this is a failure mode that most posts like this one skip, which is that evidence has a cost. Every snapshot, captured action, and logged prompt is something a reviewer has to read or deliberately ignore. Collect enough of it and you have built a haystack, with the one thing the hiring team needed sitting somewhere inside it under twenty minutes of scrubbing.
Two things break when evidence turns into volume. The reviewer stops reading, so the richest report gets the same tired glance as the thinnest one. And the candidate starts to feel watched instead of assessed, which changes how they work in ways that make the signal worse. Someone performing for a log writes different code than someone solving a problem.
So the discipline is to capture what the candidate acted on. A prompt that led to a kept change and a caught bug is signal. A raw stream of every keypress is noise wearing a lab coat. The proof a reviewer actually wants is short: what was the goal, what ran, where did the candidate look, and what did they do when something failed. Evidence that answers none of those has no reason to be in the report.
The scarce resource here is reviewer attention, and a report that spends it carelessly is worse than a shorter one that respects it.
The report has two audiences
Most scoring conversations assume one reader, the hiring team. An AI-native report has a second reader it cannot afford to ignore, which is the candidate. That second reader is easy to design away, because they are not in the room when the report gets built. Designing for them anyway is the difference between a tool that ranks people and one they would trust to have ranked them.
If a report is legible only to the people making the decision, it is a verdict, not an assessment. The candidate should be able to open the same five competencies and see why a clean technical submission still scored low on verification, or why accepting everything unread cost them on prompting even though the tests passed. Readability there keeps the scoring honest, because a report you can explain to the person it judged is one you have actually reasoned through, and a report you cannot explain is usually one the tooling decided for you.
It also pulls candidate behavior in the right direction. When people can see that reading a diff and catching a failure is the thing being rewarded, they read diffs and catch failures. The report stops being a hidden rubric to game and starts describing the skill the job needs.
What this does not claim
This does not make scoring final, magic, or fully automatic. Review still matters, and so do calibration and human judgment on the calls a model cannot settle.
The narrower and more useful point is that AI-native assessments need artifacts that make judgment inspectable. If the product asks a candidate to build in a modern workflow, the report should show how they used that workflow.
