Shared Agent Workspaces Need Evidence, Not Ceremony
Assessment agents got more useful when they stopped living in a separate file universe. A field note on shared workspaces, attribution, validation, and reviewer trust.
The first version of an assessment agent is usually impressive for the wrong reason. It feels like a second workspace appears beside the candidate: a private agent lane, a private set of files, a special promote button, and a transcript that only partly explains what changed. That is dramatic, but it is not the shape hiring teams need. In an assessment, the important question is whether the final artifact can explain who did what, when the work changed, and how the result was checked.
That pushed us toward a simpler rule: agents are chats, and chats work on the same workspace. The candidate, editor chat, and assessment agents all operate against one set of files. When an agent suggests a change, it lands through the same review path as any other AI edit. When the candidate keeps or rejects it, that decision belongs to the evidence trail. The product becomes less theatrical, but the review gets stronger.
Why separate workspaces fail
A separate agent workspace creates ceremony. You have to promote work, reconcile files, and explain why the agent's copy differs from the editor's copy. It also creates a subtle trust problem. If the reviewer sees a polished final file but cannot tell whether the candidate wrote it, accepted it, or merely inherited it from an agent run, the assessment is measuring the wrong thing.
Shared workspaces remove that ambiguity by making AI assistance inspectable. The transcript, diff, attribution, and validation proof can all refer to the same artifact because there is only one artifact.
Sharing one set of files carries a cost we had to design against. An agent that rewrites a module the candidate was mid-thought on can erase the very context a reviewer needs, and knowing which line the agent wrote does not tell you whether it overwrote a better idea underneath. The answer is to make agent edits land as proposals the candidate resolves, so a change to shared state always arrives with a human decision beside it. Handing the agent its own sandbox again would drop us right back into the ceremony we left.
Attribution is product infrastructure
Once agents share the workspace, attribution becomes part of what gets scored. A line can be written by the candidate, by the editor chat, or by an agent. A candidate can read a diff before keeping it, keep it blindly, reject it after inspection, or ask a follow-up. Those actions should not be flattened into a vague "used AI" flag.
This is the distinction AI-native assessment needs. Two candidates can both use an agent and show very different skill. One directs the agent, inspects the diff, tests the behavior, and explains the tradeoff. Another accepts every generated change without reading. How to Assess Developers Who Code With AI follows two such sessions through the same task.
Validation has to be visible
We also tightened how validation shows up, because browser checks, replays, plan cards, and walkthroughs only help if a reviewer can see what was checked. That is why validation is better as an evidence surface than as a hidden gate. AI-Native Assessments Need Evidence, Not Theater covers what a useful proof has to say.
The candidate should be able to run or trigger checks. The agent can request validation when the work calls for it. The reviewer should see enough of the path to distinguish a tight verify loop from a last-second batch check.
Seeing enough of the path does not mean seeing every step of it, because a trail that records every keystroke and agent turn can bury the two moments that mattered. It has to be ranked, with the rejected diff and the caught failure first, as the post on assessing developers who code with AI shows.
Evidence has to outlast the session
Most assessment review happens after the candidate has closed the tab. Nobody is there to explain a choice, walk back a mistake, or narrate why a diff got rejected. A validation proof that only makes sense while the session is live is a memory that fades.
So the trail has to stand on its own. A reviewer picking it up cold should be able to reconstruct the shape of the work: what the goal was, where the candidate pushed back on the agent, which check caught which failure, and what the final decision rested on. A recruiter, a hiring manager, and a skeptical engineer reading three weeks later should reach the same read of the same session, because the artifact carries its own explanation instead of leaning on the room it was made in.
The product lesson
The lesson we took from agent-native editors is to make the agent's work feel native to the workspace. For AlgoArena, that means fewer parallel universes and more durable artifacts. If a candidate uses AI well, the product should preserve the shape of that skill. If a candidate over-delegates, the product should make that visible without turning the review into guesswork.
The assessment surface is still evolving, but the direction is clear. Shared files, attributed edits, readable diffs, validation proofs, and replayable context belong together. That is how an AI-assisted session becomes something a hiring team can evaluate instead of a demo they have to trust on vibes.
