An AI fluency rubric for interviews: score the answer rather than the candidate
An interview hands you a candidate’s account of their AI work, and never the work itself. That’s the whole problem with grading AI fluency in an hour, and a rubric that stays quiet about the gap ends up carrying more weight than it earned.
Zapier made the skill concrete. It published four levels, put them to work across its hiring process, and in March 2026 it raised the bar: Capable now means AI is embedded in core work, used through repeatable systems and linked to a clear improvement in quality, efficiency or another relevant outcome. Zapier gets more than one look at the same person because those levels run across a whole process. You’re applying them to one story.
So the thing an interview can grade is the account. A candidate is describing work you didn’t see, at a company you can’t audit, and you have no idea what access they were given. The level you write down says how much evidence that account carried and how well it held up once you pushed. Everything below is built around that limit.
What an interview can actually establish
Two people with the same title can work with AI completely differently. One opens a chat window when a first draft is due. The other has built it into research, analysis or operations and knows the three places a human still has to look. Both of them will tell you they’re good with AI, and you have forty minutes to work out which is which.
Why teams started asking
The motive teams state out loud is defensive. A team that’s already rebuilt its own work around these tools is protecting the way it operates.
We started asking candidates in the hiring process what their use of AI is. It’s more to not have the risk of hiring someone who would be AI-averse.”
The question is also new. Across a corpus of 5.5 million captured conversations, a topic taxonomy tagged something as AI-fluency probing in 14.98% of the 331,872 sessions where that taxonomy fired at all. A separate measure follows one normalized topic label, AI Usage, quarter by quarter. It sat at 0.33% of interviews in the third quarter of 2025 and 4.54% in the second quarter of 2026.¹
Those two figures measure two different things. They sit on different denominators and different definitions, and both are proxy detections from a topic taxonomy that nobody has checked against human labelling, so each undercounts by an amount nobody has measured. Put together they support something narrow: the question gets asked more often than it used to, and most interviews still don’t ask it at all.
None of that adds up to a standard. There’s no settled bar for this competency and no agreed read of what a good answer looks like, and that’s exactly the condition under which a rubric gets over-trusted. Metaview’s data covers what happens inside the hiring process. It holds no post-hire outcome data at all, nothing about how anyone did once they had the job. So a rating from this rubric has nothing to be validated against, and it should never reach a hiring manager as a forecast.
Further reading: Zapier on raising its hiring bar and Metaview on AI fluency in hiring.
What the levels do and do not measure
Here’s the working definition: AI fluency is using AI on real work while staying accountable for what comes out. The version below adapts Zapier’s published levels for one interview. No company’s internal hiring rubric is reproduced here, and the level a role needs depends on the role.
Read the table as four descriptions of an answer, one per level, and never as four descriptions of a person. The third column is an illustration written for this article to show the shape such an answer takes. No candidate said any of it, and none of it comes from a transcript.
| Level | What the answer has to contain | Illustration, written for this article | What the level cannot tell you |
|---|---|---|---|
| 1. Unacceptable | No repeatable use, or no example the candidate can take past the tool name. | Names three tools, then cannot say what any of them changed about the work or how the output was checked. | Whether the candidate cannot work this way, was never given access, made a defensible choice to avoid it, or simply explained it badly. Four findings, one label. |
| 2. Capable | One recurring workflow described end to end: why that tool, what the check was, and a before-and-after the candidate states plainly. | Walks through a weekly task, names the step where the output was wrong, names who or what caught it. | Whether the workflow is as routine as described. One story is one story. |
| 3. Adoptive | Evidence that other people used something the candidate built, and that it changed after they saw it fail. | Describes a shared process, how weak output came to light, and what was rewritten in response. | How much of this was permission. Building for other people requires an employer that lets you. |
| 4. Transformative | A before-and-after at team or function level: what stopped, what replaced it, and where a human still holds the decision. | Says what the team no longer does, what it does instead, and who signs off on the output now. | Anything at all about a candidate who has never been given that scope. Reserve it for roles that will grant it. |
Where the ladder breaks
The four levels don’t measure one thing at rising strength. The first two are about a person’s own work. The top two ask for reach across other people: something built that colleagues used, a function reorganised. That’s partly a skill and partly a record of what a previous employer allowed. A contractor under a client NDA, an engineer in a regulated bank, anyone whose last team banned the tools outright: they’re all capped near the bottom of the ladder however well they work.
Unacceptable has the same problem in reverse. It bundles four separate findings under one word: couldn’t, wasn’t allowed to, chose not to for a reason you’d respect, and couldn’t explain it. Only the last two are about the candidate, and the last one might just mean the person explained it badly. Write the reason next to the rating every time, or the rating gets read as the first of the four.
What a testable answer contains
An account made entirely of wins can’t be tested, because there’s nothing in it to check against. Put a failure in it, along with the check that caught the failure and the change that followed, and now you have three specific things to push on. That’s why the failure question does more work than any other one you’ll ask.
AI is biased toward positivity. If you’re building a candidate template, you have to be prescriptive about flags.”
The same instinct applies on your side of the table. If nobody asks for the flags, an enthusiastic candidate and an enthusiastic model both hand over the positives and stop there.
Questions that test the account
Each question below closes one gap in the evidence. None of them establishes that the candidate did good work. What they establish is whether the story survives contact with a follow-up, which is a smaller claim and the only one an hour can carry. The same discipline runs through any attempt to interview for AI judgment.
1. Ask for one workflow, start to finish
Ask: “Walk me through the last meaningful piece of work where you used AI, from start to finish.” This one tests specificity. Real work comes with boring detail: the file that was wrong, the colleague who reviewed it, the second attempt.
- What outcome were you after?
- What did you hand the tool, and why that tool?
- What changed about the quality, the speed or the scope?
- What did you check, change or throw away before you used the output?
Don’t read the level off the tool name. A candidate who names the newest model has told you what they’ve opened and nothing about what they’ve shipped.
2. Ask what it got wrong
Ask: “Tell me about a time AI gave you a convincing answer that was wrong. How did you catch it?” This tests whether the check they described a minute ago is a habit or just a sentence. Listen for a named mechanism: a source somebody verified, a test case, a comparison against data the candidate already trusted, a colleague who read it over. And listen for who owns the mistake.
3. Ask what changed afterwards
Ask: “Have you changed the way you or your team work since then?” This separates a repeatable workflow from a good afternoon. A candidate who adjusted the prompt, added a review step or dropped the tool for that task is describing something that’s run more than once.
4. Ask where they stopped
Ask: “Tell me about a task where you decided not to use AI, or limited how you used it. What risk were you managing?” This tests judgment about risk, and it’s the one question a candidate from a locked-down environment can answer as well as anyone else. Useful answers name the risk:
- Confidential or personal information
- Security or legal restrictions
- Source material nobody had verified
- Bias or unfair outcomes
- Decisions that need a human on the record
- Tasks where the tool adds work
“I never use AI” isn’t an answer, and it isn’t a rejection either. Ask what the constraint was and who set it.
Give the competency one owner
Give AI fluency to one interviewer and let them ask all four questions. Four people asking “do you use AI?” in four rooms will get you four shallow accounts, and then the panel averages those into a level nobody can defend.
Writing the rating down
A rating that claims to be about the evidence has to point at the evidence. That’s a documentation problem before it’s a judgment problem. Write the scorecard two days later and you keep the impression while losing the detail that separated Capable from Adoptive: which check, whose review, what the before-and-after actually was.
Metaview Notetaker records the interview, with consent, so the candidate’s own words are still there when the scorecard gets written. Metaview drafts the scorecard from that conversation, then the interviewer edits it, sets the rating and submits it. Metaview doesn’t score the candidate and doesn’t decide anything.
None of that establishes that the account was true. A recording shows what a candidate said, and no recruiting stack can confirm they did it. The only things that test the claim itself sit outside the interview: a work sample built on the same kind of task, a paid trial, a reference who watched the work happen. Every one of those costs more than an hour, which is exactly why the rubric keeps getting asked to do a job it can’t do.
What the record does support is a better argument between humans. Metaview Reports let you query your own interview data for two things worth knowing:
- Coverage: was AI fluency raised in the stage where the plan said it would be?
- Disagreement: where did two interviewers rate comparable answers differently?
Neither one tells you which of the two was right. A disagreement is where the calibration conversation starts, and the transcript is what keeps that conversation on the answer instead of on whose memory is better. Related: Metaview’s 2026 AI & Hiring Alignment Report.
Calibrating the panel
A shared reference point still leaves you with unshared scoring. Nothing here shows that two interviewers reading one answer will land on the same level, and nothing in the data behind this article tests inter-rater agreement either. Consistency between interviewers is measurable once you have the recordings, so measure it yourself before these levels carry weight in a decision.
- Rate the same answer twice: take one recorded response, have two or three interviewers rate it independently, and compare.
- Argue the split: find the sentence each interviewer was weighing wherever the ratings differ.
- Write down what moves a rating: if nobody can say what evidence would have made it one level higher, the boundary isn’t defined yet.
- Set the required level as policy: Capable for roles where AI belongs in the person’s own work, Adoptive for roles expected to build something colleagues use, Transformative only where the job will actually grant that scope.
Leave AI out of the interview plan when it isn’t material to the role. Add a competency because it’s current, hand it to people who haven’t agreed what the levels mean, and what you’ve added to the decision is noise.
The awkward part doesn’t go away. A candidate who tells a tidy story about modest work will out-rate a stronger practitioner who explains badly, and this rubric doesn’t fix that. What it does is make the disagreement visible: the same words, two levels, one conversation about which sentence carried the evidence. Reach for a work sample when that isn’t enough for the role, and don’t bolt on a fifth level.
Run the rubric inside your own interviews.
Metaview records the interview with consent, so a debrief argues about the words that were said instead of about who remembers them better.
Frequently asked questions
What does an AI fluency rubric measure in an interview?
It measures the account a candidate gives of their own AI work: how specific that account is, whether it contains a check the candidate actually ran, and whether it holds up under follow-up questions. It doesn’t measure the work, and nothing about it forecasts how the person will do the job.
How do you interview for AI fluency?
Take one recent piece of work and stay on it. Ask what the candidate handed to the tool and why, what came back wrong, how they caught it, what they changed afterwards, and where they decided not to use AI at all. A thin account runs out of detail under those follow-ups, which is the only thing an hour can reliably test.
What are the four levels?
They’re adapted from the levels Zapier published. Unacceptable means the account showed no repeatable, meaningful AI use. Capable means AI sits in core work with checks and a stated before-and-after. Adoptive means other people used something the candidate built. Transformative means a team or function was rebuilt around it.
Can a candidate rate low because a previous employer never let them use AI?
Yes, and that’s the biggest weakness in the ladder. The top two levels ask for evidence of reach across other people, which depends on the access and the permission a previous employer granted. Write down the reason for a low rating alongside the rating, so the panel can tell inability apart from a locked-down environment.
How does Metaview help?
Metaview Notetaker records the interview with consent, so a rating can be checked against the candidate’s own words instead of a two-day-old memory, and Metaview Reports let you query your own interview data for where the competency was covered and where two raters disagreed. Neither one establishes that the account was true.
¹ Metaview corpus of 5.5 million captured conversations (2026). The 14.98% figure counts sessions where the topic taxonomy fired at all (n=331,872). The 0.33% and 4.54% figures follow one normalized topic label, AI Usage, by quarter. Both are proxy detections that have not been checked against human labelling, so read them as a floor under how often the question came up. Zapier: AI fluency levels (2 February 2026) and the raised hiring bar (31 March 2026).

