The AI-Augmented IC: How to Score Real Output When AI Did Half the Work
A take-home used to tell you something. Someone spent a weekend on it, and what came back was a fair proxy for how they think, because you couldn’t produce it without thinking. That proxy is gone. Give the same take-home to anyone with a model and a couple of hours and it comes back clean, structured, and hard to tell apart from the work of someone who couldn’t have written a line of it alone.
The same goes for most of what an individual contributor hands in. A designer’s portfolio, an analyst’s sample project, an engineer’s tidy pull request, a marketer’s writing sample: each used to be evidence, because building it took the capability you were hiring for. Now a candidate can steer a model to any of them and understand almost none of what came out.
If you’re doing the hiring, that’s a scoring problem. Using AI is the job now, so the question you’re answering is harder than it used to be: when the output is this good and this fast, how much of the judgment behind it is the candidate’s own?
Why the finished artifact stopped proving the person
Metaview’s 2026 AI & Hiring Alignment Report, surveying 505 recruiting leaders and hiring managers across North America and EMEA, found that 85% of companies that exceed their hiring goals use AI in hiring. AI is now how the work gets done, so it’s also how candidates produce the very work you’re using to judge them.
So the old evidence stops holding. Producing the work and understanding it have come apart. Someone can ship a strong artifact with a shaky grasp of it, and someone with real judgment can ship the same artifact in a tenth of the time. From the outside, on the page, the two look identical.
You were always hiring for what sits underneath the artifact anyway: can this person make good calls once the work gets messy? Marisa Uranga Bradwell, Director of Recruiting Ops at Deliveroo, keeps the focus where it belongs:
It's not just the hours saved, but also the time you can use to focus on what really matters: getting the signal we need to make the best hiring decisions.”
The evidence has moved into what a candidate can tell you about the work.
The four levels of AI-augmented output
Your panel needs the same picture of what the levels look like before it can score AI-augmented work the same way twice. The rubric below asks one question of the work: how much of the judgment behind it belongs to the candidate. It runs from output they can’t account for to output they improved on, and it works the same for engineering, design, analysis, or writing.
| Level | What you see in the walkthrough | What it tells you about the candidate |
|---|---|---|
| 1. Can’t rebuild it | Describes what the artifact does but not why it is built that way. Stalls when you ask them to change one assumption. Treats the model’s output as a finished thing they collected. | The output is the model’s, and so is the judgment. Fine for throwaway work, a real risk anywhere the reasoning has to be theirs. |
| 2. Can explain it, didn’t shape it | Understands the output and can walk you through it, but every meaningful choice was the model’s default. Took what it produced and tidied the edges. | Competent, and easy to overrate. They can operate the tool, but you have not watched them make a hard call yet. |
| 3. Directed it | Set the constraints, caught where the model went wrong, and made real tradeoffs. Can point to the parts that are theirs and the parts that are the model’s. | The judgment is there. Strong for most roles, because the work gets better when they are the one steering. |
| 4. Out-judged it | Used the model to move fast and still overrode it where its default was wrong for this context. Can name the better path and why the obvious one fails here. | The profile worth holding out for. They get the speed of AI and keep the judgment the model does not have. |
Most candidates land at Level 2 or Level 3, and from a finished artifact alone you can’t tell the two apart. The interview exists to do exactly that. Level 4 is rarer, and it turns up in the same place.
How to score it: run the walkthrough
You can’t read judgment off a document. You get it by watching someone reason about their own work in real time, with no model there to fill the gaps. Three moves cover it, and none of them needs a new take-home.
Make them rebuild one decision live
Pick one non-obvious choice in what they submitted and ask them to rebuild it from scratch. ‘Why this approach here, and not the simpler one?’ Whoever made that call answers in seconds. Whoever took the model’s default has to reverse-engineer a reason while you watch, and you can hear it happening.
Follow the reasoning behind the work
The thinking around the work is what you’re grading now. Ask what they tried and threw away, where the model was confidently wrong, and what they’d change given another day.
Strong candidates keep a running commentary on their own work. Weaker ones can only defend the surface of it, and that shows up fast once you push on a detail.
Probe the road they didn’t take
Ask about the option they rejected. ‘What was the obvious alternative here, and why didn’t you go with it?’
Most of someone’s judgment sits in the options they turned down. That’s the hardest part to fake, because the reasoning behind it never made it into the artifact.
The evidence is in the texture of the answer, and that texture is the first thing you lose by the time you sit down to score. So keep the conversation on the record.
Metaview’s interview notes captures every spoken word, so the moment a candidate reconstructed a tradeoff, or couldn’t, is sitting there to score against two days later instead of being pieced back together from memory.
Score from what was said
A great artifact throws a halo, and that halo is what makes judgment hard to score fairly. The candidate who handed in the cleanest project feels like the strongest hire, even when the walkthrough said otherwise. The fix is to score from evidence, on the same rubric, for everyone. With the interview captured, you can point to the exact moment a candidate reasoned through a tradeoff, or reached for the model instead, and settle on a level with the evidence in front of the room.
A rubric only counts for something when the whole panel applies it the same way. Matthias Schmeisser, Global Senior Director of TA at emnify, describes the shift once the scoring runs on structure:
I've fallen in love with features like Candidate Comparison which allows me to compare candidates based on our interview structure, and also with having an AI assistant that deep dives on the scorecard. It's just next level.”
Build it into how your team hires
None of this needs a new stage. Swap the graded take-home for a short walkthrough of work the candidate already did, score it on the four-level rubric, and put judgment under AI on the scorecard next to communication and ownership.
Use interview notes to capture the walkthrough, and reports to show which questions your interviewers actually covered across the pipeline. It works the same way when an agent takes the first pass on sourcing: you judge the shortlist with its reasoning attached.
Want the groundwork first? We’ve written up quality of hire and what separates a great interviewer, and the Alignment Report has the data behind the shift.
Every hiring team is working through this right now, and the answer is more hopeful than the panic suggests:
AI closed the gap on producing the artifact. The judgment behind it is still all the candidate’s, and a good interview is where you get to see it. Score for that and you’re hiring for what matters more every time the models get better.
See the work your candidates really did.
Metaview captures the conversation, structures it against your rubric, and lets your panel score AI-augmented work from what a candidate actually said, so a polished output never gets the benefit of the doubt.
Frequently asked questions
What if a candidate did not use AI on the work you are reviewing?
The walkthrough works the same way. Judgment shows up whether the first draft came from a model or from a blank page, so you're still asking whether they can account for the decisions in the work and defend the calls they made. Someone who did the thinking answers with the same fluency no matter what produced the draft, and the tool they reached for matters far less than whether the reasoning is theirs.
Does scoring AI-augmented work differ for engineers, designers, and writers?
The method stays the same across roles; what changes is the decision you pick to walk through. For an engineer that might be a data-model tradeoff; for a designer, a layout or flow choice; for a writer, a structural call about what to cut. You run the same three moves and let the candidate reconstruct a decision that's specific to their craft.
What if the candidate has nothing they can walk you through?
Give them a small, realistic task to do with a model while you watch, then run the same three moves over what they just produced. You're scoring the reasoning as it happens, so a fresh piece works as well as something they brought. It also removes any question of who really did the work.
How much AI use is too much on a take-home?
There's no ceiling worth setting. Someone can lean on a model for most of the drafting and still show more judgment than a candidate who typed every line by hand, because the score comes from the calls they can defend under questioning. Set the task so it's only strong when the human made good decisions, then read for those decisions.
How do you get a new interviewer scoring on the four-level scale?
Have them score a recorded walkthrough on their own before they see how the panel scored it. Where their level differs from everyone else's is exactly where their read of judgment needs calibrating, and you can go back to the specific answer to talk it through. Two or three rounds of that and a new interviewer lands in the same place as the rest of the panel.