How to interview for working with AI agents
Teams are adding a line about AI to the hiring bar faster than they’re adding a question about it to the interview. Shopify is the best known example. It told its people that reflexive AI use is now a baseline expectation, and one company saying so still isn’t a trend anyone has counted. You can count what happens in the interview afterwards, and mostly the line gets waved through on a confident answer and a tool name.
AI fluency or AI tool use came up in 14.98% of the 331,872 interview sessions where the topics discussed could be detected. The subject is climbing fast from a low base: AI usage as a discussed topic went from 0.33% of conversations in the third quarter of 2025 to 4.54% in the second quarter of 2026. Both figures come from topic detection across Metaview’s aggregate corpus of 5.2 million candidate interviews, and topic detection is a proxy. It tells you the subject came up. It can’t tell you how well anyone assessed it.
A rubric won’t close that distance on its own, so it’s worth being blunt about what one can do. It can’t tell you who will be good with agents once they start, because nobody holds that evidence and Metaview holds no data on what happens after someone is hired. A rubric turns a vague line on a scorecard into a defined one, so four interviewers assess the same thing and their disagreement lands on the candidate instead of on what the words meant. That’s a smaller claim than most rubric articles make, and it’s the part that holds up.
Why the line keeps getting skipped
The interview is already full. That, more than indifference, is why the question keeps getting skipped. A single interview in that corpus touches a mean of 17.52 detected competencies, so a new one tends to arrive as the last two minutes of a loop, asked by whoever remembers. Adding it properly means deciding what it is and what it displaces, and that conversation runs slower than writing a phrase into a scorecard. The phrase spreads faster than the practice.
There’s a reasonable case that the trouble is warranted, and it lands better from someone who hires for a living than from a vendor.
If the muscle is atrophied, it will enhance an atrophied muscle. If the muscle is strong, it will enhance the strong muscle.”
That’s a practitioner’s read, and nobody has measured it. It argues for interviewing carefully, and it stops well short of showing that any particular level means anything. Judgment is the thing being assessed if an agent tracks the judgment of the person holding it, and judgment has never been readable from a tool name.
What a rubric can and cannot tell you
Start with the limit, because everything else depends on it. A delegation rubric scores an account of past work. The work itself never enters the room. A candidate who describes how they scoped a task, what context they gave, and where the output drifted is handing you a description, and a description can be detailed and rehearsed at the same time. Specificity gives you better material to probe, and it doesn’t make the account true.
The second limit: nothing connects a level to what comes after the offer. Metaview’s corpus covers what took place inside the hiring process and holds nothing about what happened once people started, and the four levels below haven’t been validated against anything at all. Read a level as a description of what the candidate told you, and leave the predictions out of it.
You still get something worth having. Four interviewers who share a set of anchors produce answers you can lay next to each other. A level with the evidence attached still reads a week later, at the debrief, when nobody can quote the example any more. The disagreement gets a subject when two people land on different levels for the same candidate: which part of the account they read differently, and what would settle it. A vague bar produces vague arguments, and the loudest voice in the room tends to win those.
- Whether the account is accurate, rehearsed, or borrowed from a teammate.
- Whether the person can do the same thing on your stack, under your constraints.
- How they will work once they start. Nobody holds that data, including us.
- Which level is objectively better. The right level is the one the role needs.
- What the candidate described scoping, handing over, and checking.
- That four interviewers used the same anchors, so the calls can be compared.
- Which evidence sat behind the call, still readable at the debrief.
- Exactly what two interviewers disagree about when they disagree.
The rubric is built against one specific trap, the one that shows up whenever a competency is new and nobody has agreed what good looks like.
Hiring managers conflate activity with progress.”
The interview rewards whoever described the most AI use when nobody has set anchors, because volume of activity is the only thing an unanchored answer can be compared on. Anchors give the room something else to compare.
The four-level delegation rubric
Here’s one way to cut it. Four levels is a choice, and there’s nothing canonical about the number or the names. They earn their place only if your team agrees on them before the loop starts and uses the same words in the debrief. Read the fourth column hardest. It holds what the level doesn’t establish, which is most of what a level is.
| Level | What the candidate describes | What you tend to hear | What this level is not evidence of |
|---|---|---|---|
| 1. Operator | Does the work themselves. Uses AI for small pieces at most, and rarely hands over a whole task. | “I tried it, but it was faster to do it myself.” Names a tool they stopped using. | Weak judgment. Someone who decided a task was too consequential to hand over has made a defensible call, and a rubric that punishes it is measuring enthusiasm. |
| 2. Delegator | Hands discrete, well-defined tasks to an agent and takes the output mostly as it comes back. | “I had it write the first version, then I cleaned it up.” One-off tasks, light checking. | A ceiling. Plenty of work needs nothing more, and the account may only describe what the last job gave them room to do. |
| 3. Supervisor | Scopes multi-step work, sets the context and the limits, reviews what comes back, and can say where it tends to break. | “I gave it the brief and examples, caught where it drifted, and changed how I set it up after that.” | That the checking was any good. You are hearing that they checked, which is not the same as knowing what they caught or what they missed. |
| 4. Orchestrator | Runs several agents at once, and has built repeatable workflows that other people use. | “I built the workflow the rest of the team runs on.” | Ability on its own. Building something a team adopts also takes access, authority and time, and plenty of jobs hand out none of those. |
The gap teams argue about most sits between delegator and supervisor, and it comes down to whether the person treats the agent as confident and often wrong. Define that gap carefully, because it’s where two interviewers most often score the same answer differently. Calling it a validated threshold would be inventing a finding.
The levels only sort anything if the questions surface behavior instead of opinions. Three moves do most of that work:
- Ask for one handoff. “Walk me through the last real piece of work you handed to an agent, end to end. What did you hand over, and what did you keep?” A specific example gives you something to probe, and it says nothing about whether the work went the way it’s being described, which is why the follow-ups matter more than the story.
- Inspect the limits they set. How did they scope it, what context did they give, and what did they check before using what came back? A delegator tends to stop at “it gave me a draft.” A supervisor can narrate how they kept hold of it.
- Probe the catch. “When did an agent do something confidently wrong, and how did you notice?” This is where a thin answer usually shows, because the detail that makes a failure story credible is hard to improvise.
Write down what a strong, a middling and a weak answer sounds like at each level before the loop starts, and give the question to one interviewer instead of all of them. Spread it across a panel and you get four people asking whether the candidate uses AI, with nobody going deep enough to place them anywhere.
Where the disagreement should happen
All of this assumes the answer is still there when the score gets written. A delegation answer runs long and specific, and the specifics fade first. The anchors you wrote turn into decoration once the level gets assigned from an impression, and the debrief is back to comparing confidence.
That’s what capture is for. The Notetaker joins the interview as a participant with consent and records the conversation, then turns it into structured notes, each section linked back to the moment in the transcript it came from. Metaview drafts the scorecard from that conversation, and the interviewer reviews, rates and submits it. The level stays the interviewer’s call, made against the anchors the team wrote, from an answer that’s still on the page.
One level up, Reports covers whether a pipeline assessed the competency at all, and how the calls break down by team. That’s how you find out that half a loop skipped the question instead of scoring it low, and where two interviewers sit consistently a level apart. Closing that gap is a conversation your team has to have. A report shows you the divergence. It can’t calibrate anyone.
One piece of context belongs here, and the caveat goes first. Metaview’s 2026 AI and Hiring Alignment Report asked recruiting teams about their own use of AI. It asked nothing about the candidates they hired, so none of it is evidence about a delegation rubric, and it’s a single cross-sectional read, so what it found are associations.
The survey covered 505 respondents, 252 recruiting leaders and 253 hiring managers at companies with 200 or more employees across North America and EMEA. 85% of the companies that had already exceeded their hiring goals use AI in hiring at all. Teams where AI is core to hiring rate the recruiter and hiring manager relationship as excellent 55% of the time, and the report puts that at 3.8x the rate among teams that don’t use AI. 79% of all 505 respondents say they’re optimistic about AI’s future in hiring. Those are facts about recruiting teams, and the question here is about candidates. They’re part of why anyone is asking it at all, and no part of the answer.
What this means for your team
Pick one role you’re hiring for now and settle, out loud with the hiring manager, which level it actually needs. Most roles don’t need the top of the table. Writing that level down, with two lines on what a strong answer sounds like, is most of the work, and it’s the part that keeps getting skipped.
Then put it where the interview happens. Give the question to one interviewer, add it to your question bank and scorecard templates, and keep the answer available for the debrief. Revisit the anchors after a quarter of loops, because the first version of any rubric is wrong in ways only four people using it can show you. Our writeup on what separates good interviewers from bad ones covers the groundwork underneath all of this.
Working with agents landed on the hiring bar faster than most processes adapted, and honestly, nobody yet knows what any given interview answer means once the person starts. You can define the thing you’re asking about, ask it the same way every time, and keep the evidence so the argument at the debrief is about the candidate. That’s less than a rubric usually promises, and it’s the part that will still be true in a year.
Interview for it, and keep the answer
Capture the conversation, place the candidate against your own anchors, and keep the evidence readable when the debrief happens.
Frequently asked questions
What does it mean to interview for working with AI agents?
It means assessing how a candidate describes handing real work to an AI agent: how they scoped the task, what context and limits they set, what they checked before using the output, and how they noticed when the agent was confidently wrong. It’s closer to a delegation question than a tool question, so you interview for it the way you would for delegation and judgment rather than for whether someone has used a chatbot.
What are the levels of working with AI agents?
Four levels cover most roles. An operator does the work themselves and rarely hands it over. A delegator hands discrete tasks to an agent and takes the output mostly as it comes back. A supervisor scopes multi-step work, sets limits, checks the result, and can say where the agent tends to break. An orchestrator runs several agents and has built workflows other people use. The four are a shared vocabulary rather than a validated scale, and the right level is the one the role needs.
How do you interview for AI delegation skills?
Ask the candidate to walk you through one real piece of work they handed to an AI agent, end to end. Then ask how they scoped it, what context they gave, and what they checked before using what came back, and ask about a time the agent was confidently wrong and how they noticed. A specific account gives you something to probe. It doesn’t confirm that the work happened the way it was described, which is why the follow-up questions carry more weight than the story itself.
Does a delegation rubric predict how well someone will work with AI once they are hired?
No. A rubric scores an account given in an interview, and nothing links that score to what happens after someone starts. Metaview holds no data on what happens after a hire, and these levels haven’t been validated against anything. Use a level to make four interviewers comparable, and to give the debrief something specific to argue about.
How does Metaview help assess working with AI agents?
With consent, the Notetaker joins the interview and records the conversation, then turns it into structured notes linked back to the transcript, so the candidate’s example is still available when the scorecard gets written. Metaview drafts the scorecard from that conversation, and the interviewer reviews, rates and submits it. Reporting shows whether the competency was assessed across the loop. The level itself stays a human call.