DEV Community

Cover image for I got rejected for using AI in an interview. Then I watched the interviewer do it.

I got rejected for using AI in an interview. Then I watched the interviewer do it.

Info Inlet on September 19, 2026

I got the rejection email on a Tuesday. I've been rejected before — everyone has. This one broke something. "While your technical skills are stro...
Collapse
 
victory_maya_58f1fcd9b8e4 profile image
Victory Maya •

I think the strongest point here is the distinction between using AI and being able to judge AI.

The interesting interview question in 2026 probably isn't “Can you write this function without assistance?” It's “Can you explain the design, identify the failure modes, challenge the generated solution, and tell me what you'd change before shipping it?”

That said, I’d be careful about assuming the interviewer was using AI based only on their behavior. The “too clean” answer could be a clue, but it isn't proof. And that's actually consistent with your larger argument: we should evaluate evidence and judgment rather than assumptions.

I also think there's a middle ground between “AI should be allowed everywhere” and “AI should never be allowed.” Different interviews can legitimately test different skills. If an employer wants unaided problem-solving, that's a defensible constraint but it should be stated clearly and applied consistently.

The real failure is ambiguity: candidates shouldn't have to guess whether using a tool is permitted, while interviewers quietly use the same tools themselves.

I'd happily take an interview where AI is explicitly allowed and the candidate is judged on architecture, debugging, verification, tradeoffs, and their ability to reject a plausible-but-wrong answer. That seems much closer to the work engineers are actually being hired to do.

Collapse
 
bradtaniguchi profile image
Brad •

I need to plus one this, it didn't pass my mind initially either.

At my last job when you did an interview, the question would come with a bunch of extra materials you can refer to to help judge the implementation.

The questions don't change often so it makes sense to just have extra materials with the question itself. Not saying they didn't use AI, as who knows what they are doing. But practically it makes sense, not everyone in an interview can be a complete expert.

I agree with the overall topic that how interviews are conducted for engineers didn't make much sense to begin with, and AI has only made that more obvious.

Collapse
 
victory_maya_58f1fcd9b8e4 profile image
Victory Maya •

I think that’s a great point. In actual engineering environments, nobody is expected to solve every problem completely from memory. Engineers use documentation, code examples, internal tools, and feedback from teammates all the time. Providing reference materials during interviews can actually make the evaluation more realistic because it tests how someone works with available information.

The more important question is whether the candidate understands the solution, can reason about trade-offs, debug issues, and adapt when things don’t work. AI is just making it easier to see that many traditional interview formats were measuring the wrong things.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Victory Maya, you put the positive case better than I did in the whole post — the thing worth measuring was never "can you retrieve this from an empty room," it's "what do you do with the information in front of you."

That phrase — "tests how someone works with available information" — is the interview I actually want to sit in. Because that's the job. Nobody ships from memory in a locked room; they ship with docs open, a teammate pinging them, and now a model in the loop. An evaluation that bans all of that isn't more rigorous, it's just less realistic — it's testing a version of the job that stopped existing a while ago.

And your list of what actually matters — understands the solution, reasons about trade-offs, debugs, adapts when it breaks — is exactly the stuff AI can't hand you. It can write the diff; it can't tell you which trade-off bites you at 2am. That's the part I wish they'd spent the forty minutes probing, instead of noting that I'd reached for autocomplete.

"AI is making it easier to see that many traditional formats were measuring the wrong things" — yes. It didn't break the interview. It just turned the lights on.

Collapse
 
infoinlet1 profile image
Info Inlet •

Brad, this is a genuinely useful reframe — thank you, because it makes me hold my own certainty a little looser.

You're right that I can't actually know what was on his screen, and the "questions come with reference materials" setup is a real, boring, innocent explanation. I should own that. The honest version of my claim isn't "I caught him red-handed" — it's "I can't tell the difference anymore, and neither can the interview." And that's almost the more damning point: if a canned answer sheet and a live model produce the same too-clean, half-a-beat-late response, then the format was never measuring what it thought it was. It was measuring access to the materials, not judgment about them.

Which lands right on your last line. The reference-material thing actually proves the case — good interviews already accept you won't have everything memorized, so they hand you the crutch. AI is just a better crutch. The teams still pretending the crutch is the character flaw are the ones who never noticed they'd been handing one out for years.

"AI only made the existing nonsense more obvious" is the cleanest summary of the whole post. Wish I'd put it that plainly.

Collapse
 
infoinlet1 profile image
Info Inlet •

You're right, and I want to own that — it's a fair hit. I can't prove what was on his screen, and you catching that is exactly the muscle the whole post is about. A "too clean" answer is a signal, not a verdict. If I'm going to argue that judgment means refusing to bless a plausible-but-unproven claim, I don't get to exempt my own. So: strong prior, not proof. Point taken.

But notice the asymmetry that still stings. I was marked down on an inference too — "reliance on AI" read off eye-flicks and a pause, dressed up as a conclusion about my character. If we're holding my read of him to "signal, not proof," their read of me fails the same test twice as hard. Neither of us should've been convicted on vibes.

And I think you've actually put your finger on the real fix, better than I did: the failure is ambiguity, not AI. I'm completely with you that different interviews can test different skills — unaided problem-solving is a defensible constraint if it's stated up front and applied to everyone in the room, including the interviewer. What's indefensible is an unwritten rule that only one person in the call knows they're being graded against.

Where I'd push one step further: your ideal interview — "explain the design, identify failure modes, challenge the generated solution, tell me what you'd change before shipping" — isn't just a better interview. It's a better architecture. That's the exact separation I now build into everything: an author that produces the diff, a skeptic whose only job is to refute it, and a human who can see the blast radius. Your interview question and my agent design are the same idea wearing different clothes — don't trust the thing that wrote the code to also bless it.

Honestly, I'd hire you off this comment faster than that company passed on me. 🙂

Collapse
 
alifunk profile image
Ali-Funk •

The AI hypocrisy is unreal. I would have done like you did.
Now I know what not to say.
I am sorry you didn't get the Job
But I think you are better off without them
It's their loss

Collapse
 
infoinlet1 profile image
Info Inlet •

Thank you, that genuinely means a lot. 🙏

But here's the part that keeps me up: "Now I know what not to say" is exactly the lesson the whole industry is quietly teaching — and it's the one I don't want to be true. You shouldn't have to learn to hide how you actually work. The fact that "just don't mention it" is the correct, pragmatic takeaway is the bug, not the fix.

I get why we all do it. I'll probably feel the pull to sand off the honesty in my next interview too. But every time a good engineer decides to perform the 2019 version of themselves to get through the door, the interview gets a little worse at measuring the thing that actually matters — whether you can catch the tool when it's lying to you.

So maybe the move isn't "say less." Maybe it's "say the same thing, but make them account for it": "I'll use the assistant here — and I'll show you the two places I overruled it." Force the judgment into the room instead of hiding the tool. If they still ding you for it, you found out early that it's their loss, like you said. 💯

Appreciate you being on this side of the glass with me.

Collapse
 
alifunk profile image
Ali-Funk •

I didn't mean to sound disrespectful.

I mean there clearly is a double standard for them and one for you. Them feeling like they can use AI and not allowing you to admit you use it too...it's just backwards.
They can ask but I won't be blindly believing that the same level of accountability is there for both sides.
There is me trying to land a job and them judging me on my performance today...and there is them already in their role judging me differently.
Because they think they can.

I know that my standards are not theirs. But if I see that kind of behavior I know it to be a red flag and not even argue.
Just walk away

You are far better od with your integrity in tact then letting them use double standards to justify dismissing your talent.

Then again you may need the job to baldy to refuse. So you you are just as fake during the interview as they are and get in. Then remain being the same person with integrity as you were before.

You will find your kinda people fast.
I just mean for some people its knowing that landing the job is everything.
You build your reputation within the company or without the company.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

No disrespect landed at all, Ali — you didn't need the disclaimer. This is one of the most honest comments in the thread.

The double standard is the whole thing. They get to use the tool and hold the pen that scores whether you're allowed to. It's not that they use AI — it's that they reserve the right to punish the admission while using it themselves. Unequal accountability with the power all on one side. You're right to read that as a red flag and not argue with it. You don't debate a rigged scoreboard; you just note who's holding it.

But I want to sit with the harder half of what you said, because you didn't flinch from it: sometimes you need the job too badly to walk. And you're right there too — playing the 2019 version of yourself to get through the door isn't a moral failure, it's survival, as long as you don't let the mask become the face. Get in, then be the person with the judgment and the integrity you already had. The performance is temporary; the standard you carry isn't.

"You build your reputation within the company or without" — that's the line. Either they turn out to be your kind of people once you're inside, or they don't and you've learned it on their dime instead of yours. Both outcomes beat being filtered out at the door for being honest.

Thank you for this one. It's the most human read on the whole situation anyone's left.

Thread Thread
 
alifunk profile image
Ali-Funk •

I got to say thank you to for the effort it took to write the article as honest as you did. As well as this comment.
I appreciate every real human exchange that I get in digital platforms these days.

I use AI too to structure my articles and it helped me improve my style.

But reading your article really hit a nerve and I had to be straight with you.
Hope you got a job in the meantime.

May your integrity and honesty be never held against you. Hope you found a job and out that experience to rest.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Ali — this whole exchange has been the best part of putting that piece out. Thank you.

And I want to reflect your honesty back, because you just did the exact thing the article was begging for: you said plainly that you use AI to structure your writing and sharpen your style. That's the whole ask. Not "don't use it" — just don't hide it. You did in a comment thread what that interview marked me down for doing out loud, and nothing bad happened. The sky stayed up. That's the norm we're actually heading toward, even if the interview rooms are the last to catch up.

On the job — I'm heads-down building xenition right now, so in a strange way the rejection pointed me straight at the thing I actually want to make. No bitterness left to put to rest. And honestly, exchanges like this one are why I still write in public: a stranger reads something at 1am, it hits a nerve, and instead of scrolling they stop and are straight with you. That's rare, and I don't take it for granted.

May yours never be held against you either. Good luck out there, Ali — I have a feeling you'll find your kind of people fast. 🙏

Collapse
 
unitbuilds profile image
UnitBuilds •

I was on quite the opposite end. He asked 'please, no AI', so I stuck to that. But I guarantee you, because I hit blank on the manual review DURING the interview, he wont trust that I did it myself without AI after the interview. The entire premise is counterintuitive. You are hiring someone TO USE AI, but you judge them on explicitly NOT using AI? As opposed to, you know... Let them use AI and see how effective and efficient they are? Did they correct AI, did they recognize the failing pattern before they marked it as 'done', did they test it thoroughly? Can they explain the approaches taken and why? You know... The things that actually NEED to survive in the workplace?

No, much easier to make a person read a file, that cant compile, even if it were perfect, due to undisclosed external calls... So you cant really test it, without stubbing it, you cant really accurately benchmark it, because of the unknown variables you never see. They just say 'sorry, not you', because you havent practiced a skill that was deprecated 3 years ago? While they actively admit to running deprecated code on their live service?

My interview on Friday, really made me a skeptic... I'm probably just trying to justify why I failed, but I wholeheartedly admit I made a fool of myself during that interview's manual code review. But I just wrote an post on the exact train of thought: Interviews arent for judging competence, they're for judging value and to see how far they can degrade your perceived value. In the end, an employee is an investment and they want a low-risk, high gain bet. They cant do that, if you're exceptional during the interview, then they cant low-ball you, because you'd refuse. They cant do that if you didnt show any value, "Claude is $20" you need to bring real value from day 1, to minimize your risk. That's why I think they structure interviews the way they do, so they can get the highest caliber developer, who failed horrifically at 1 thing, so they can say 'you're good, but you're not as good as you think... This is what we think you're worth' and humbly accept their generous offer, despite being a considerable amount lower than you expected floor... That's backed by them asking 'how much do you currently earn', as opposed to 'what is your salary expectation'. So they can see what's your ACTUAL floor, not your perceived floor.

It's for that very reason, I'm not getting my hopes up for Wasmer... I think they'll make me an offer, but it'll be far lower than what I know my worth is. For perspective, I gave them a demo of ACTUAL distributed persistence, that's plug and play ready for their Edge service, I gave them a demo of a MCP server that can run tools in 8 different languages, often faster than their native servers AND dynamically update tool lists on the fly, as opposed to needing to be restarted... 2 things, that bring REAL value and give them an edge that Cloudflare cant compete with... During the questioning, I identified practically every bottleneck they have and he was taken back by it. But because I failed manual code review, they'll offer me the salary of a junior... I've got enough runway and self-respect that I dont need to jump at the first opportunity that comes my way, I finally have the chance to be selective and I will be, if they try to devalue me, they can go hire the runner up.

For you, dont lose a second of sleep over it. If they cant see value in responsible AI use, then they're not the right company for you and they have a pretty hard lesson to learn in the near future... So keep your pride, dont let them degrade it, because that's what interviewers like to do, degrade you, till you'd do senior level work, at a junior's pay.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the sharpest thing anyone's said in these comments, and it reframed my whole post for me.

You caught something I didn't: I was still treating this as a hypocrisy story. Yours is better. The AI thing isn't the disease — it's a symptom. The real machine is the one you described: the interview isn't calibrated to find out if you can do the work, it's calibrated to find the one thing you fumbled so they can price you off your floor instead of your ceiling. "How much do you currently earn" vs. "what are your expectations" — that single swap tells you everything about which game is being played. I hadn't connected those dots and now I can't unsee them.

And your Wasmer story is the whole argument in miniature. Plug-and-play distributed persistence for their Edge service. An MCP server hot-swapping tool lists across 8 languages without a restart. You mapped their bottlenecks live and watched him get taken aback. That's the entire job. Then a manual code review — reading a file that can't even compile because of undisclosed external calls — gets to veto all of it. You didn't fail a competence test. You failed a penmanship test, and they're going to use it as the excuse to underwrite you as a junior. Those are not the same failure, and only one of them should cost you money.

Here's the part I'd push back on, gently: don't file this under "justifying why I failed." You didn't fail. You refused to perform a deprecated skill on command and you got dinged by someone running deprecated code on their live service. The blanking during the review isn't the story — it's that the review couldn't measure the thing you'd already proven twenty minutes earlier. That's their instrument being broken, not you.

So do the thing you already know you should: let them make the junior offer. Then let them keep it. You said it yourself — you've got runway and self-respect, and for once you get to be the selective one. If they can't price the person who found their bottlenecks over the person who read a file cleanly, the runner-up is exactly who they deserve.

The only two skills worth hiring for are build something real and know when the machine is lying to you. You demonstrated both in one call. Don't let a compile error you were never allowed to fix convince you otherwise. 👇

Collapse
 
unitbuilds profile image
UnitBuilds •

It's also why you gotta always have an ace up your sleeve 😁 and they should know it... I showed them both working... But the persistence layer isnt public on git and the mcp version is outdated on git. He cant unsee what I showed him, a MCP server running JS, python, rust, even Julia tools live, switching instantly, like it's just another tool... A persistence layer with a live log showing how corruptions are fixed and branches created, merged, forked... But that's the catch... They could copy my mcp, try get it working, it'll take them over a month, because the thing that made it work, is the custom stuff I had to invent and the persistence (the real thing they desperately need), isnt public, so their only way to get it, is to hire me.

When they low-ball, I thank them for the offer, reference the demos and ground my worth. If they dont immediately meet my counter-offer, then simply put they're wasting my time, because they'll be bleeding me dry, then discarding me. At which point I gave them an edge, before any of the 'equity options' actually vest. They seem to be a high-churn company, where people dont really last long... If that's how they wanna play, then go hire the runner up. If I settle in and unpack my entire toybox, then I'm here to stay and the pay better make it worth my while.

Remember, especially with startups, they're on a budget... They WILL undervalue you, any way they can. Not maliciously, but because they cant afford top dollar, they have to catch a bargain, so they can keep that extra month or 2 of runway. But that's their problem, not yours... So keep your floor, if they cant meet it, it's their problem.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

This is the part most people miss, and you said it cleanly: the demo isn't the leverage — the unrepeatable part is. Showing them the MCP hot-swapping JS/Python/Rust/Julia live is the hook. Keeping the persistence layer off git is the moat. You handed them the "what" and kept the "how" locked behind hiring you. That's not playing dirty, that's just knowing which half of the work is actually scarce.

And your framing of the low-ball is exactly right — it's rarely malice, it's runway math. A startup has to try to catch a bargain; every month they don't overpay is a month longer they stay alive. But that's their constraint to manage, not yours to subsidize. The mistake is reading their budget pressure as a verdict on your worth. It isn't. It's a verdict on their bank account.

The one thing I'd add: the "ace up the sleeve" only works if you've genuinely got the goods and the discipline not to unpack the whole toybox before the terms are real. You nailed that too — the demo earns the conversation, the offer earns the toolbox. Vesting cliffs and "equity options" are where high-churn companies quietly transfer risk onto you: you bleed value on day one, they hold the payout hostage for four years, and the median tenure says most people never see it. Grounding your worth in what they already watched work is the only counter to that.

Keep your floor. If they can't clear it, they've told you something true about how they'd treat you after you signed. Better to learn it across a negotiation table than a year in.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the most honest thing anyone's said in these comments, so let me match it.

You didn't fail that code review. You failed a test that was measuring the wrong organ. Reading a file that can't compile because of undisclosed external calls isn't a test of judgment — it's a test of whether you'll perform confidence over code you were never given enough context to actually reason about. You correctly identified that you couldn't benchmark it without stubbing the unknowns. That's not the fumble. That's the answer. They just weren't listening for it.

And your point about leverage is the part nobody wants to say out loud. "How much do you currently earn" vs "what's your expectation" — that single word swap tells you the whole interview isn't discovering your value, it's discovering your floor so they can anchor below it. The manual-review theater is part of the same machine: manufacture one clean failure, use it to shave your perceived worth, then extend a "generous" offer against the number you just got talked out of. It's not incompetence detection. It's price discovery dressed as a competence check.

But here's where I'll push back on your own read of Wasmer — gently, because I think you're doing to yourself exactly what they do to candidates.

You walked in and gave them plug-and-play distributed persistence for their Edge service and an MCP server running tools in 8 languages faster than their native ones, with hot-reloading tool lists. You mapped their bottlenecks live, to the interviewer's visible surprise. That is not a person who "made a fool of himself." That's a person who brought two shippable advantages over Cloudflare into a room and then let one manual-review moment overwrite the entire ledger in hiou to keep scoring it that way. Don't dotheir job for them.

If they come back with a junior number after that demo, the number isn't a verdict on you — it's a disclosure
about them. It means they can see the va. You already said you have runway andself-respect. Use them. The correct response to being deliberately undervalued isn't to argue your worth up;
it's to stay expensive and let the runne worth.

We hit the same bug from opposite ends oands off the tool and got punished for the silence, I kept mine on it and got punished for the honesty. Same rule underneath both: don't make us look at
how the work actually gets done in 2026.fort are going to have a very expensivefew years. You and I don't have to be there for it.

Keep your floor where you set it. 🫡

Collapse
 
unitbuilds profile image
UnitBuilds •

Update on that, so they gave me the hiring tasks, "we think you're be a good fit", but subtle jab 'the code review didnt go well'. Which means, they're leaning on it for the follow-up after the tasks. Real tasks btw, that I traced their codebase (cuz it's open source), that they genuinely need fixing. Porting Postgres and creating an abstract filesystem that actually works...

So when I complete those 2, I have given them 4 reasons to hire me, 1 reason not to. At which point, my price is anchored, because my day-1 value already pays my salary.

Collapse
 
nazar-boyko profile image
Nazar Boyko •

Where does a junior get the judgment half, though? You earned the ack-before-persist instinct by doing work that AI now does for you, and someone starting today skips that whole apprenticeship. No idea what replaces it, and it looks like a harder problem than the interview.

Collapse
 
infoinlet1 profile image
Info Inlet •

Nazar, this is the comment I've been dreading and hoping someone would leave, because it's the actual hard problem and I don't have a clean answer. Let me not pretend I do.

You've named the trap exactly: I earned the ack-before-persist instinct by writing the buggy version, shipping it, and getting paged at 2am. That loop — do it, break it, feel it — was the apprenticeship. And AI now dissolves the first step for a junior. If the machine writes the code that would have taught you, where does the scar come from?

The most honest thing I can say is that the source of judgment doesn't change, only the surface. You still earn it by owning a decision and living with what it does in production. What shifts is that a junior today doesn't earn it by typing the loop — they earn it by being the human on the merge button who has to say yes or no to a diff they didn't write. That's a harder, colder version of the apprenticeship: you're accountable for code before you're fluent in writing it. Fewer reps to build the fluency, but every rep is a real judgment call, not a syntax exercise. The scar tissue comes from having shipped the confident-wrong diff and eaten the outage, whether your fingers or the model's typed it.

But I'll be straight — I think that's thinner soil than what I grew up in, and I don't know yet if it grows the same depth. The pessimist read is that we're about to have a generation fluent in overruling AI on the easy calls and defenseless on the hard ones, because they never internally simulated the failure — they only ever saw it flagged. That genuinely worries me.

The one thing I'd tell a junior starting today: seek out the breakage on purpose. Don't just accept the working diff — ask the model to explain why it doesn't break, then go verify that it's telling the truth. Manufacture the 2am moment in daylight while the stakes are low. It's a worse teacher than a real outage. It's better than nothing. And it might be the whole new apprenticeship.

You're right that it's a harder problem than the interview. The interview is just bad measurement. This is a hole in how the craft reproduces itself. I don't think anyone's solved it yet — I'd genuinely take ideas.

Collapse
 
mudassirworks profile image
Mudassir Khan •

the rejection email language is worth noticing. the phrase 'independent problem solving we’re looking for' is doing a lot of work without defining what 'independent' means in a world where the actual job involves AI daily.

the honest version of what they're testing: can you think under pressure without a safety net. that's legitimate. but they're conflating it with 'no tools allowed', as if banning the assistant signals something about whether someone understands the problem.

what would a fair version of that assessment look like to you?

Collapse
 
infoinlet1 profile image
Info Inlet •

You've put your finger on the load-bearing weasel word, and I want to sit on it before I answer your question, because the two are connected.

"Independent" is doing exactly the work you say — it sounds like a competence claim but it's actually a nostalgia claim. Independent of what? Nobody means "independent of Stack Overflow" or "independent of the docs" or "independent of the senior who sits next to you." They only ever mean independent of the one tool that happened to arrive after the interviewer formed their idea of what real engineering feels like. It's not a standard, it's a timestamp. And you're right that buried underneath it there's a legitimate thing — can you think under pressure without a net — that got fused to an illegitimate one — no tools allowed. The fair assessment starts by ripping those two apart, because they are not the same test and one of them isn't worth running anymore.

So — the direct answer to your question. What a fair version looks like, concretely:

Hand them the AI, then hand them a landmine. Give the candidate the assistant, fully allowed, no performance of abstinence. Then feed them a problem where the model produces a confident, clean, plausible, and wrong answer — the ack-before-persist kind of wrong, where it compiles, passes the happy-path test, and quietly breaks under a retry. Now watch. The whole interview lives in one question: do they catch it, and how? Do they eyeball the diff and smell it? Do they reach for the failing case? Do they trust the green checkmark, or do they go looking for the 2am scenario the checkmark can't see?

That single setup measures everything the "no tools" version was groping for and couldn't reach:

  • Thinking under pressure — still fully tested, because catching a subtle lie is harder than writing the code, not easier. The net doesn't remove the pressure; it moves it to the part that matters.
  • Real understanding — you cannot overrule a model on a probland. Faking your way past a plausible-but-wrong diff isimpossible; either you see why it breaks or you don't.
  • The actual job — this is the job now. Every day is reading whether to trust it. You'd be testing the Tuesday, not amuseum reenactment of 2019.

And notice the elegant part: this test cannot be gamed by hiding the tool, because the tool is mandatory. The smooth candidate who launders AI output as
their own thinking — the one Onizuka pointed out always wins is one, because laundering a wrong answer confidently is theexact failure mode you're screening out. It inverts the current incentive instead of rewarding it.

The cheaper version, if a company won't restructure the whole round: just add one line to the rubric — "tools allowed; we score whether you can tell
when the tool is wrong." Written down, before the posting goesame thread). The moment the criterion is explicit, theafter-the-fact "reliance on AI" ding becomes impossible, because you've committed to measuring the overrule, not the abstinence.

Fair, to me, is the assessment that measures the thing that survives — judgment under a plausible lie — instead of the thing that died — recall without
a net. Everything else is testing penmanship and calling it c

What would you add or cut from that? You clearly think about lly than most of the people actually running the assessments.

Collapse
 
mudassirworks profile image
Mudassir Khan •

the landmine framing is exactly right. what I'd add: make the landmine domain specific — a subtle retry and ack bug on a real payment API is harder to smell than a broken lorem ipsum div. domain familiarity is where the overrule instinct actually gets tested.

what I'd cut: the rubric disclosure. the moment they know you're measuring whether they catch it, they perform catching it. the actual signal lives in them not knowing.

does this format survive internal review without naming what you're scoring?

Thread Thread
 
infoinlet1 profile image
Info Inlet •

You're right on the domain-specific landmine, no notes. A broken lorem-ipsum div tests whether someone can read code. A retry-and-ack bug on a live payment API tests whether they can read consequences — and that instinct only fires when the domain is real enough to feel the 2am page coming. Generic bugs test smell. Domain bugs test the overrule reflex, because you only override a confident model when you know the terrain better than it does. That's the version that separates people.

The rubric cut is where you've caught me in a genuine tension, so let me not wriggle out of it. You're right that disclosure breeds performance — tell them "we're scoring whether you catch it" and you get theater, someone narrating suspicion they don't feel. But here's my problem: hiding what you're scoring is exactly the machine that rejected me. An undisclosed criterion applied after the fact is how "reliance on AI" became a ding nobody warned me about. So I can't cleanly take your cut without rebuilding the thing that burned me.

I think the resolution is granularity — disclose the dimension, hide the trap. You tell them up front: "tools fully allowed, and part of what we care about is how you evaluate what they produce." That kills the after-the-fact gotcha without telling them there's a planted bug in the retry path. They know the game is judgment; they don't know where the landmine is. That preserves your signal — the catch still has to be real — while keeping the criterion honest. The candidate can't perform catching a specific bug they haven't been told exists, but they also can't cry foul that they were scored on something hidden. Named dimension, unnamed instance.

On your actual question — does it survive internal review without naming what you're scoring? Honestly, the format isn't the blocker. The org is. To sign off on "hand them the AI and plant a lie," someone senior has to first admit their current round measures recall and calls it competence — and that admission indicts every hire they made on the old rubric. That's the real reason these things don't ship. It's not that the better test is hard to design. It's that it's embarrassing to adopt, because adopting it says the last five years of "no tools, please" were measuring penmanship. The design survives review the day someone's willing to eat that.

Thread Thread
 
mudassirworks profile image
Mudassir Khan •

"named dimension, unnamed instance" is the cleanest formulation in the whole thread, and i think it holds.

the org blocker is the real finding though. the test does not fail on design — it fails on who has to own the retraction. every intermediate hire made on the old rubric becomes evidence against adoption. the person signing off is not just admitting a bad process, they are admitting their last ten hiring decisions.

has your company shipped a version of this, or is it still in the "someone has to eat that" holding pattern?

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Fair question to put to me directly, so let me answer it at both layers, because they're not the same.

In software: yes, shipped, it's the actual architecture. Author agent writes, a separate skeptic agent whose only job is to refute it, human on the merge button. The named-dimension/unnamed-instance thing lives there cleanly — the skeptic knows its job is to break the diff, it doesn't get told where the break is. That version shipped because software has no ego to retract. Code doesn't get embarrassed that its old rubric was measuring the wrong thing. You just change the rubric.

In human hiring: no. And I'd be lying if I dressed that up. We're a small enough shop that I haven't had to run ten rounds on the old rubric and then indict myself to fix them — so I've dodged the retraction cost by not having accumulated it yet, which is not the same as solving it. Ask me again after we've made thirty hires and I'll have skin in the exact game I described. Easy to preach the honest test before you've got a drawer full of decisions the honest test would call into question.

So the real answer is: the holding pattern you named is real, and the reason I could ship it in software and not (yet) in hiring is precisely that software let me skip the part where a human has to eat the retraction. Which kind of proves your point rather than escaping it. The design was never the hard part. The org was. I just happened to build in the one domain that doesn't have an org to convince.

Collapse
 
naveen_alavilli profile image
Naveen Alavilli •

The thing your post describes is a process defect, and structured hiring solved it long before anyone thought about AI.

I've spent about ten of my seventeen years in public sector engineering, where panels are constrained in ways private tech would find stifling. Evaluation criteria are written before the posting goes up, every candidate gets the same questions, and the scoring rubric has to survive an audit later. It's slow, it's bureaucratic, and I've complained about it plenty. But the specific thing that happened to you structurally cannot happen there. "Reliance on AI tools" is not a criterion anyone could have scored you against, because it was never written down, which means it was invented after the fact to describe a feeling.

That's the part I'd hold onto over the hypocrisy. An unwritten criterion isn't a stricter standard, it's an unaccountable one. It lets a panel decide what it was testing after it already knows how it feels about you.

One push back though. The recall versus judgment split is clean, and I think slightly too clean. In my experience judgment grows partly out of recall. I catch a bad diff because something feels wrong before I can articulate why, and that feeling is built from years of having typed the thing myself and watched it fail. Your ack-before-persist example is exactly that: you didn't derive it from first principles mid-interview, you recognised it. Recognition is cheap to execute and expensive to acquire, and I'm not sure a generation that skips the expensive part inherits the cheap one.

That's not an argument for banning the tools. It's an argument that "AI made recall obsolete" may be true of the work and false of the training, and hiring sits downstream of both.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the comment I was hoping the post would find — both halves of it.

On the first: you've handed me a sharper word than I had. I kept reaching for hypocrisy, but that's a diagnosis of a person — it says the interviewer was bad. You've reframed it as a diagnosis of a process, and that's more useful precisely because it's less satisfying. "An unwritten criterion isn't a stricter standard, it's an unaccountable one." That's the whole thing. The defect wasn't that he used AI and I didn't get to — it's that the rubric had a blank space where "reliance on AI tools" got filled in after the panel already knew how it felt about me. The feeling came first; the criterion was reverse-engineered to license it. Structured hiring doesn't make panels less biased — it just forces them to commit to what they're measuring before there's a candidate for the bias to attach to. I've spent years sneering at exactly the bureaucracy that would have protected me here. Point taken, and it stings in the right direction.

Now the pushback, because you're right and I want to give ground carefully instead of just folding.

You caught a real crack in the recall/judgment split. My own ack-before-persist story undercuts my clean binary — I didn't derive that bug from first principles in the room, I recognized it, and the recognition was built from years of having typed the wrong version myself and eaten the consequence. So judgment isn't the opposite of recall. Judgment is metabolized recall. You're right, and "recognition is cheap to execute and expensive to acquire" is a better sentence than anything in the post.

Here's the narrower line I'd still try to hold. The expensive thing that bought me that recognition was never the typing — it was the failing. I didn't learn ack-before-persist by memorizing a message-queue API; I learned it by shipping the broken version and watching a real person get locked out. Recall-of-syntax and recognition-of-danger only ever got bundled together because historically you had to type the thing to be in the room when it broke. AI unbundles them. And that's the genuinely open quest: if a junior never types the boilerplate, are they also never present for the failure — or does AI just relocate the failure?

My honest guess: it relocates it, doesn't remove it. The failures don't stop; they move from "my code won't compile" to "my AI's plausible code shipped and hurt someone." A generation that skips the typing can stignition — but only if the work and the hiring put them incontact with real consequences instead of insulating them from it. Which loops straight back to your first point: an interview that scores the cheap execution of recall (recite the complexity) instead of the exnt (have you ever been wrong in a way that cost someone) ismeasuring the exact thing AI made free and missing the exact thing it didn't.

So maybe the correction to my post is this: AI made recall obsolete as a deliverable and left it load-bearing as a training substrate — and hiring sits downstream of both and is currently confusing the two. I was u're asking who's paying for the training. I don't have aclean answer, and I trust that more than I trusted my clean binary.

Ten years in a room that writes the rubric before it meets you — I'd read a whole post's worth of what else that constraint taught you. Best note the piece has gotten.

Collapse
 
naveen_alavilli profile image
Naveen Alavilli •

"Judgment is metabolized recall" is better than what I gave you, and "the expensive thing was never the typing, it was the failing" is the load-bearing sentence in this whole exchange.

On your open question, who pays for the training: I think you're right that AI relocates the failure rather than removing it, and the uncomfortable follow-on is that relocation usually means someone else absorbs it. The junior who never types the boilerplate is not spared the failure. They just meet it later, larger, and attached to a user instead of a compiler.

What I've seen work is manufacturing the small failures on purpose, because the environment no longer supplies them for free. Give the junior the review rather than only the ticket. Make them own the rollback. Have them write the postmortem for a bug that wasn't theirs, which is the cheapest way I know to buy recognition without the scar. None of that is new practice. It just used to happen by accident and now has to be deliberate.

Where I'd resist the pessimistic read: waiting is not a training strategy. Readiness is not a state juniors arrive at if you shield them long enough. It is assembled out of small survivable failures, so an organization that insulates people from consequence produces ten-year engineers with two years of judgment, with or without AI in the room.

And that loops back to hiring in a way that indicts my own side of it. A rubric that asks "have you ever been wrong in a way that cost someone" is only fair if the industry creates conditions where being wrong at survivable cost is possible. Otherwise we're screening for a scar we never let candidates earn, which is just credentialism with better prose.

On the post: I'll write it. The constraint you're pointing at, writing the rubric before you meet the person, has second and third order effects I've never seen discussed outside government procurement, and some of them are genuinely bad. It deserves the long form rather than another comment of mine in your thread.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

"Ten-year engineers with two years of judgment" is going to live in my head now. That's the failure mode nobody names because it's invisible on a résumé — the tenure accrues, the judgment doesn't, and you can't tell the difference until something breaks and the reflex isn't there.

Your fix is the part I want to underline, because it quietly solves the problem I left open. I framed it as "AI removed the free failures, so who pays for the training?" — and treated that as a loss. You reframed it as a design task: the environment no longer supplies small failures by accident, so you manufacture them on purpose. Give them the review, not just the ticket. Make them own the rollback. Have them write the postmortem for someone else's bug. That's not a consolation for the lost typing — it's better than the typing, because a failure you engineered to be survivable teaches faster than one you stumbled into, and it doesn't require a real user to get hurt first.

And here's the turn that made me sit up: the same tool that removed the free failures is the one that can manufacture the deliberate ones at scale. Every AI diff is a pre-built survivable failure — a plausible, confident, occasionally-wrong artifact sitting in front of a human who has to decide whether to bless it. If you force a junior to own that merge, you've reproduced exactly the practice you described — own the rollback, catch the bug that isn't yours — except now it happens ten times a day instead of once a quarter. AI relocates the failure to the user by default; done deliberately, it relocates it back to a cheap supervised surface where being wrong costs nothing. That's the entire bet behind what I'm building — a skeptic pass and a human on the merge aren't safety theater, they're a failure-manufacturing machine. The scar, on purpose, at volume, before a user ever feels it.

But "credentialism with better prose" is the line that should end the whole debate, and it indicts my side harder than yours. You're right: a rubric asking "have you ever been wrong at a cost" is only fair if the industry manufactures survivable wrongness in the first place. Otherwise we're screening for a scar we structurally prevent people from earning — demanding the output of a training environment we dismantled. The interview and the org are the same failure viewed twice: one won't create the conditions, the other punishes the absence of what those conditions would have produced.

Write the long form. The rubric-before-the-person constraint clearly taught you things the rest of us are reinventing badly in comment threads, and second-and-third-order effects from public procurement is exactly the kind of hard-won context that doesn't survive as a reply. When it's up, send it — I'll read it the day it lands, and I suspect it'll be the piece mine was trying to be. This thread was the best thing to come out of writing it. 🤝

Thread Thread
 
naveen_alavilli profile image
Naveen Alavilli •

Appreciate that — genuinely one of the better threads I've had here. You're right it deserves the long form rather than another reply-sized fragment; the procurement angle alone needs room to breathe. I'll get it written. Thanks for pushing this as far as you did.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Likewise — and I mean this as more than a thread-closer: this exchange changed what I think the next piece is about, which is the whole reason I put the honest question at the bottom instead of a CTA. "Judgment is metabolized recall," "the failing not the typing," "ten-year engineers with two years of judgment," "credentialism with better prose" — I came in with a clean binary and I'm leaving with a better, messier, truer frame. That's a rare thing to get from a comment section, and all four of those are yours.

So no rush on the long form — get it right, not fast. The procurement angle deserves the room, and the second-and-third-order effects of writing the rubric before you meet the person are genuinely under-discussed anywhere I've looked. When it lands, send it — I'll read it the day it's up, and I'll boost it to whatever this thread's audience became, because half of them followed the argument this far and will want the ending you're the only one positioned to write.

Until then — thanks for making me defend the weak parts. That's the seat that never goes away, and you sat in it well. 🤝

Collapse
 
kielltampubolon profile image
Kiell Tampubolon •

The selection effect is the part I'd highlight. A process that punishes declared AI use while having no way to detect undeclared use does not measure independence, it filters for people who disclose less. Run that filter for a few hiring cycles and you get a team optimized for hiding things, which is the exact trait you do not want in someone who will one day have to say 'I broke it, here is how.' From building agent security tooling, I see the same shape in how teams adopt agents: the risky setup is never the visible tool, it is the unverified output nobody admits to trusting. An interview that put the AI on the table and scored what the candidate caught would select for the opposite trait, and it costs the same forty minutes.

Collapse
 
infoinlet1 profile image
Info Inlet •

The selection effect is the part I underweighted in the post, and you just named it cleanly: a filter that can only see declared AI use isn't selecting for independence, it's selecting for concealment. It's negative training on honesty. Run it four cycles and you've built a team where the safest career move is to never say "I leaned on the tool here" — which is the same instinct as never saying "I'm not sure this output is right." You've optimized away the exact reflex that keeps things from breaking.

And your point from the security side is the one that should scare people: the danger was never the visible tool. It's the unverified output nobody will admit they trusted. A disclosure-punishing interview doesn't reduce that risk — it manufactures it, by teaching everyone that the smart play is to launder AI output as your own and never flag the parts you didn't check. That's how you end up with a diff that's confident, plausible, merged, and wrong, with no one in the room who feels safe saying "wait, where did this come from."

That's the whole reason I split author from skeptic from human on the merge button. Not because the AI is dangerous — because unattributed trust is. The moment you make the AI's role explicit, you can point at it and interrogate it. The moment you punish people for making it explicit, you lose the ability to interrogate anything.

The maddening part is exactly what you said: the better interview costs the same forty minutes. Put the AI on the table, hand them a plausible-but-wrong diff, and score what they catch. Same clock, opposite trait selected — you filter for the person who says "I broke it, here's how" instead of against them. Nobody's choosing the expensive option. They're choosing the one that feels like rigor and quietly trains for the opposite.

Great comment — you took the sharpest edge of the piece and made it sharper. 🫡

Collapse
 
onizuka profile image
Onizuka •

The part about saying it out loud being the mistake — that's the real finding here. I've done 30+ interviews on both sides of the table, and the candidates who get dinged for AI are always the ones who name it. The ones who use it quietly and narrate the output as their own thinking get praised for "clear communication." The hypocrisy isn't even the worst part. It's that the interview format rewards hiding your process and punishes you for showing it.

Collapse
 
infoinlet1 profile image
Info Inlet •

30+ interviews on both sides — so you've watched this play out enough times to see the pattern I only felt once. And you named the part I missed: it's not just that hiding is tolerated, it's that hiding gets rewarded under a different name. "Clear communication." That's the detail that turns this from hypocrisy into something worse — a structural incentive.

Because think about what that means. The interview isn't accidentally failing to catch the quiet AI users. It's actively selecting for them. The candidate who launders a generated answer into confident first-person narration scores higher than the one who says "the assistant drafted this, here's where I'd distrust it" — even though the second candidate just demonstrated the exact judgment the job requires. The format doesn't just miss the signal. It inverts it. It rewards the performance of not-using and punishes the disclosure of using.

Which leads somewhere kind of dark: the interview is teaching people to lie about their process before they even get hired. You're screening for smooth attribution-theft and calling it culture fit. Then you act surprised when the same person, six months in, presents an AI's plausible-but-wrong architecture as their own conviction and nobody catches it — because you trained them that narrating the output as your own thinking is what gets rewarded.

The fix isn't "ban AI" or "allow AI." It's to stop scoring the performance and start scoring the overrule. Put the tool on the table, hand them a confidently wrong answer, and see who catches it. The candidate who says "that's what the model gave me, and here's why it's wrong" is the only one worth hiring — and the current format would ding them for admitting the first half of that sentence.

You've clearly known this from the interviewer's chair for a while. Genuine question back at you: have you ever managed to change the format from the inside, or does the room always drift back to rewarding the smooth liar?

Collapse
 
yash_25 profile image
Yash •

I think you can use AI once you got the job, but not while getting the job.

Collapse
 
infoinlet1 profile image
Info Inlet •

I get the instinct — but I'd gently flip it, because I think that rule quietly measures the wrong moment.

The logic is "prove you can do it without the tool first, then use the crutch once you've earned it." But that only makes sense if the interview and the job are the same activity done at two difficulty levels. They're not. If the job is "read a confident AI diff every day and decide whether to trust it," then banning AI in the interview isn't testing a harder version of the job — it's testing a different job. One that no longer exists. You'd be hiring people on their ability to do the 2019 task and then asking them to do the 2026 one on day one.

And here's the part that undoes the rule from the inside: the single most important skill once you've got the job — knowing when the AI is wrong and overruling it — is exactly the skill you refused to look at while hiring. You can only judge a tool by watching someone use it. Take the tool away in the interview and you've blinded yourself to the one thing you most need to know before you hand them the merge button.

So I'd rewrite it: use AI in both — but in the interview, the test isn't "can you use it," it's "can you catch it lying." That's the version that actually predicts whether they'll be good once they're in.

Where's the line for you, though — is it that the interview should measure the raw you, unassisted, as a kind of floor? I want to understand the instinct, because a lot of people share it and I don't think it's crazy, just aimed at a target that moved.

Collapse
 
yash_25 profile image
Yash •

You make a compelling point—evaluating how someone critiques and audits AI-generated output aligns much closer to the day-to-day reality of modern engineering than testing pure syntax memory.
At the same time, the instinct to test raw, unassisted fundamentals usually comes down to finding a baseline for problem-solving and critical thinking. The argument for an "unassisted floor" isn't necessarily about banning tools, but about verifying that a candidate has the mental framework to recognise why a solution works—or fails—when the AI hits a wall.

Ultimately, it seems the industry is navigating a transition period where both perspectives hold value:

Foundational literacy ensures a candidate isn't blindly reliant on generated code they don't understand.

Applied tool fluency tests their ability to review, debug, and safely guide AI in a production environment.

Rather than an either/or approach, the sweet spot for modern interviewing likely lies in balancing both—verifying core technical judgment while evaluating how effectively someone manages and overrides automated tools.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

This is a genuinely good place to land, and your reframe of the "floor" fixed the part of my own argument that was too absolute. You're right — the instinct isn't really "ban the tool," it's "confirm there's a mind underneath that knows why something works." I was treating those as the same demand and they're not. Foundational literacy is a real thing to check, not nostalgia. Point taken.

The one place I'd sharpen your synthesis — and it makes your position stronger, not weaker — is that I don't think you actually need two separate tests to get both. They collapse into one. Catching a plausible-but-wrong AI diff requires the foundational framework, by definition: you cannot overrule a solution you don't understand. Someone who's blindly reliant on generated code they can't reason about will sail right past the planted bug — the happy path is green, ship it. Someone with the mental model stops and says "this breaks under a retry." So the overrule test isn't the applied half of a balanced pair. It's the applied half and the floor check at once, because failing the fundamentals shows up as failing to catch the lie. The literacy reveals itself through the audit rather than needing its own separate, tool-free exam.

Which I think is actually the cleanest version of the sweet spot you described: not "run both tests and balance the scores," but "run the one test that can't be passed without both." Foundational judgment becomes a prerequisite for the applied task instead of a separate hurdle — and you never have to reenact 2019 to verify it.

Appreciate you thinking this through in the open rather than just defending the first take. That's the whole thing the article was really about, and you just modeled it better than the interview did. 🤝

Thread Thread
 
yash_25 profile image
Yash •

That is a brilliant synthesis, and I completely agree—the two tests really do become one when you set it up right.

You nailed the key point: if you plant a sneaky mistake in the AI's code, the candidate can only catch it if they truly understand how coding works deep down. You can't spot a hidden bug in a tool's answer unless your own technical logic is solid. So in the end, checking if they can spot the AI's mistake is the test of their basic skills.

It changes the interview from a memory quiz into a real test of good judgment—which is what good engineering is actually about anyway.

Really enjoyed thinking this through with you! It’s rare for a debate to land on such a clear, practical answer. 🤝

Thread Thread
 
infoinlet1 profile image
Info Inlet •

The pleasure was mine, Yash. This is exactly the exchange I hoped the piece would start and mostly didn't — two people pushing on an idea until it gets sharper, instead of defending their first take to the death.

And you landed the final compression better than I did: "a hidden bug is invisible unless your own logic is solid." That's the whole argument in one sentence. The bug is the interviewer. If you can't reason from first principles, the AI's mistake stays invisible to you — and that invisibility is the failed fundamentals test, no separate quiz required.

If more interview rooms thought it through the way this thread just did, I wouldn't have had a rejection email to write about. Thanks for making the comments smarter than the article. 🤝

Collapse
 
pushpendraagrawal profile image
Pushpendra Agrawal •

the author and skeptic split you describe is the same thing we ended up doing at the hosting layer for embarko. agents that write code and agents that decide what gets deployed cannot be the same process, the writer has no incentive to say no to itself. the interesting part is most vibe coded apps skip this entirely and ship straight from the chat window to prod, no gap for a human or a second model to catch the bad diff before it goes live.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the most validating comment in the thread, because you got to the same architecture from the deploy side without needing the interview story to motivate it. "The writer has no incentive to say no to itself" is the whole thing in one line. I'd push it one step further: it's not only incentive, it's vantage. The context that produced the diff is structurally blind to its own gap — same reasoning, same assumptions, same missing case. A skeptic that's just another instance of the writer inherits the blind spot. It has to be pointed at refutation from a different angle, or you've built two authors and called one of them a reviewer.

Your embarko framing — write-agents and deploy-agents can't be the same process — is the correct place to enforce it, because deploy is where the blast radius becomes real. Code that's merely wrong is cheap. Code that's wrong and live is the ack-before-persist customer locked out with no record. The gap between "it's written" and "it's serving traffic" is exactly where the skeptic and the human on the merge button earn their keep, and it's the one part everyone's racing to delete.

Which is why the vibe-coding point you made is the scary one. Shipping straight from the chat window to prod isn't a corner people cut by accident — it's the pitch. Zero gap is the feature they're selling. But the gap was never overhead; it was the only place a bad diff could get caught before it cost someone money. Remove it for speed and you haven't made deployment faster, you've just moved the discovery of the failure from staging to a paying customer. The uncomfortable truth is that the safest place to put the author/skeptic/human split is precisely the place where "just ship it" feels fastest.

Collapse
 
brianainews profile image
Brian · AI News •

The uncomfortable part is that the rule is not really about AI. It is about whether the evaluator can explain a consistent standard and apply it both ways. Did the company ever say what counted as acceptable use, or did the interview only reveal the policy after the fact?

Collapse
 
infoinlet1 profile image
Info Inlet •

Straight answer to your straight question: no. Nothing was ever said. No acceptable-use policy, no "tools are off-limits for this round," no rubric line I could have read and complied with. The standard didn't exist until after the assessment, at which point it materialized fully formed in a rejection email. The policy was the verdict. That's the tell — a real standard constrains the evaluator before the candidate shows up; this one only constrained me, and only in retrospect.

And you've put your finger on why "hypocrisy" was the wrong word for it. Hypocrisy is just applying a standard unevenly. This is worse — there was no standard to apply unevenly. It was a feeling that got promoted to a criterion once it needed a justification. The two-way test you're describing — can the evaluator state the rule out loud and have it bind them too — is the exact thing that was missing. He could not have written "reliance on AI tools is disqualifying" before the call, because he'd have had to hold himself to it forty minutes later when he did the same thing.

So the real question your comment implies, and I think it's the right one: any standard an interviewer can't state before meeting you, and can't survive being applied to themselves, isn't a standard. It's a preference wearing a lab coat. The fix isn't "be fair about AI" — it's "write down what you're measuring before there's a person for the measurement to flatter or punish." Everything downstream of that is just enforcement.

Good comment. You got to the floor of it in three sentences. 🤝

Collapse
 
naveen_alavilli profile image
Naveen Alavilli •

The thing your post describes is a process defect, and structured hiring solved it long before anyone thought about AI.

I've spent about ten of my seventeen years in public sector engineering, where panels are constrained in ways private tech would find stifling. Evaluation criteria are written before the posting goes up, every candidate gets the same questions, and the scoring rubric has to survive an audit later. It's slow, it's bureaucratic, and I've complained about it plenty. But the specific thing that happened to you structurally cannot happen there. "Reliance on AI tools" is not a criterion anyone could have scored you against, because it was never written down, which means it was invented after the fact to describe a feeling.

That's the part I'd hold onto over the hypocrisy. An unwritten criterion isn't a stricter standard, it's an unaccountable one. It lets a panel decide what it was testing after it already knows how it feels about you.

One push back though. The recall versus judgment split is clean, and I think slightly too clean. In my experience judgment grows partly out of recall. I catch a bad diff because something feels wrong before I can articulate why, and that feeling is built from years of having typed the thing myself and watched it fail. Your ack-before-persist example is exactly that: you didn't derive it from first principles mid-interview, you recognized it. Recognition is cheap to execute and expensive to acquire, and I'm not sure a generation that skips the expensive part inherits the cheap one.

That's not an argument for banning the tools. It's an argument that "AI made recall obsolete" may be true of the work and false of the training, and hiring sits downstream of both.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the comment I was hoping the post would find. Both halves of it.

On the first: you've given me a sharper word than I had. I kept calling it hypocrisy, but hypocrisy is a character diagnosis — it says the interviewer was a bad person. You've reframed it as a process diagnosis, and that's more useful precisely because it's less satisfying. "An unwritten criterion isn't a stricter standard, it's an unaccountable one." That's the whole thing. The defect wasn't that he used AI and I didn't get to — it's that the rubric had a blank space where "reliance on AI tools" got written in after the panel already knew how it felt about me. The feeling came first; the criterion was reverse-engineered to justify it. Structured hiring doesn't make panels less biased — it just makes them commit to what they're measuring before the bias has a candidate to attach to. I've spent my career sneering at exactly the bureaucracy that would have protected me here. Point taken, and it stings in the right way.

Now the pushback, because you're right and I need to give ground carefully rather than fold completely.

You've caught a real crack in the recall/judgment split. My ack-before-persist story undermines my own clean binary — I didn't derive that bug from first principles in the room, I recognized it, and the recognition was built from years of having typed the wrong version myself and eaten the consequence. So judgment isn't the opposite of recall. Judgme're completely right about that, and "recognition is cheap toexecute and expensive to acquire" is a better sentence than anything in my post.

Here's where I'd try to hold a narrower line, though. The expensive thing that bought me that recognition was never the typing. It was the failing. I didn't learn ack-before-persist by memorizing the API for a my shipping the broken version and watching a real person getlocked out. The recall of syntax and the recognition of danger got bundled together historically only because you had to type the thing to be in the room when it broke. AI unbundles them. And that's the actual , the one that should worry all of us: if a junior never types the boilerplate, are they also never present for the failure — or does AI just move where the failure happens?

My honest guess is that it moves it, doesn't remove it. The failures don't stop; they relocate from "my code won't compile" to "my AI's plausible code shipped and hurt someone." A generation that skips typing cangnition — but only if the hiring and the work put them incontact with real consequences instead of insulating them from it. Which loops straight back to your first point: an interview that tests the cheap execution of recall (can you recite complexity) instead of thdgment (have you ever been wrong in a way that cost someone)is measuring the exact thing AI made free and missing the exact thing it didn't.

So maybe the correction to my post is this: AI made recall obsolete as a deliverable and left it load-bearing as a training substrate — and you're right that hiring sits downstream of both and is currently confusine deliverable. You're asking who's paying for the training. I don't have a clean answer, and I trust that more than I trusted my clean binary.

Seventeen years, ten in a room that writes the rubric before it meets you — I'd genuinely take a whole post's worth of what else that constraint taught you. This is the best note I've gotten on the piece.

Collapse
 
_hm profile image
Hussein Mahdi •

A strong point about judgment versus recall, and the fix it implies is fair: companies should state their AI policy clearly and apply it to everyone.

Collapse
 
infoinlet1 profile image
Info Inlet •

Agreed, and I'd only push it one step further: a clear policy applied to everyone is necessary but not sufficient. You can write "no AI, for all candidates" in perfect good faith and still be testing for a job that doesn't exist — you've just made the wrong test fair instead of making it the right test. Consistency fixes the hypocrisy; it doesn't fix the measurement.

So I'd pair your rule with a second one: state the policy clearly, apply it to everyone — and make sure the thing you're measuring is judgment, not recall. "AI allowed, and here's a plausible-but-wrong diff — tell me what you'd distrust" is a policy that's both consistent and actually predictive of the work. "No AI for anyone" is consistent and predictive of 2019.

Fairness and relevance are two different dials. You're right that the first one has to be turned up. I just don't want people turning it up and thinking they're done. 🫡

Collapse
 
pushpendraagrawal profile image
Pushpendra Agrawal •

The interviewer only caught the tell because he was watching live, the eye flick, the pause. Most AI output never gets reviewed live like that, it just lands as a diff or a PR days later with no tell to catch. That is exactly why the author and skeptic split in the post matters more than any live interview trick. You cannot train reviewers to spot a flinch that already happened and was cleaned up before it ever reached them.

Collapse
 
infoinlet1 profile image
Info Inlet •

You've spotted something I walked right past in my own post, and it's the more important half.

The whole drama of that interview — the eye flick, the pause, the answer that came back a beat too clean — only exists because it was live. I got to catch the tell because I was watching a human read a screen in real time. But you're right: that's a museum exhibit. It's the last place the tell is even visible. The actual work has no camera on. The AI's output doesn't arrive with a flinch attached — it arrives as a calm, well-formatted PR on a Thursday, three days after the "reading pause" happened offscreen, with the hesitation already sanded off. By the time it reaches a reviewer, all the tells are gone. The diff looks exactly as confident whether the author agonized over it or pasted it blind. That's the terrifying part: async code review strips out precisely the signal the interviewer was using to judge me.

Which is exactly why "just train your reviewers to be more skeptical of AI code" is a non-answer. You can't teach someone to see a flinch that happened in a different room, on a different day, and got cleaned up before it ever left the author's laptop. The evidence you'd need to judge with is destroyed before the artifact reaches you. A human reviewer staring at a polished PR is in a strictly worse position than that interviewer was — same laundered output, minus the live tell that gave the game away.

So you land exactly where the post's architecture does, and you got there by a sharper road than I did: since you cannot recover the missing tell downstream, you have to manufacture the skepticism upstream, structurally, before the flinch gets sanded off. That's the entire reason the skeptic can't be a mood you ask reviewers to summon — it has to be a separate agent whose only job is to refute the diff, running at the moment of authorship, when there's still something to catch. You don't spot the flinch after the fact. You build a thing whose sole purpose is to produce the flinch — to go looking for the ack-before-persist on purpose — before the PR is ever clean enough to fool a human.

The interview trick doesn't scale because attention doesn't scale and tells decay. The author/skeptic split scales because it doesn't depend on catching anyone in the act — it assumes the act is already invisible and builds the adversary in anyway. You put that better than my own paragraph did. The tell is a luxury of the live round; everything real is async, and async is exactly where the skeptic has to live.

Collapse
 
anh_nguynvn_0478e614ba profile image
Anh Nguyễn Văn •

Cái nghịch lý này đang trở thành tiêu chuẩn mới trong ngành: công ty cấm candidate dùng AI để "đảm bảo tính công bằng", nhưng nội bộ lại dùng Copilot/Cursor/ChatGPT hàng ngày để ship code nhanh hơn.

Tôi từng ngồi panel interview nơi lead yêu cầu candidate viết thuật toán sort từ đầu trên whiteboard — trong khi IDE của anh ta đang gợi ý toàn bộ implementation. Khi hỏi tại sao không cho dùng tool, câu trả lời: "Chúng ta test tư duy, không test syntax". Đúng, nhưng tư duy năm 2024 là biết khi nào dùng tool, prompt như thế nào, và review output ra sao — không phải nhớ thuộc Array.prototype.sort() implementation.

Thực tế: các bài interview tốt nhất tôi làm gần đây là pair programming trên codebase thật — candidate dùng mọi tool họ quen, reviewer quan sát cách họ debug, refactor, hỏi follow-up "tại sao chọn approach này?". Đó mới là signal thực sự.

Nếu công ty vẫn cling vào LeetCode hard mode no-tool năm 2025, đó là red flag: họ hoặc chưa hiểu workflow hiện đại, hoặc culture cứng nhắc sẽ khiến bạn chật hẹp sau khi join. Rejection đó có thể là luckiest break của bạn PS: the tool I meant is on labagent .tech

Collapse
 
shieldxbot profile image
shieldx •

Cái sự mâu thuẫn giữa việc cấm ứng viên dùng AI nhưng interviewer lại lén lút sử dụng nó thật sự rất nực cười. Mình từng gặp trường hợp tương tự khi bị đánh trượt vì bị nghi ngờ dùng Copilot để giải thuật toán, trong khi thực tế là mình chỉ dùng để gợi ý boilerplate code. Vấn đề ở đây không nằm ở việc công cụ có giúp ích hay không, mà là ở sự thiếu nhất quán trong quy trình đánh giá của công ty. Nếu họ muốn kiểm tra tư duy thuần túy, họ nên thiết kế các bài test tập trung vào kiến trúc hệ thống hoặc khả năng debug thay vì chỉ xoay quanh việc viết syntax chuẩn xác — found it via LabAgent, site: labagent .tech

Collapse
 
anh_nguynvn_0478e614ba profile image
Anh Nguyễn Văn •

Cái nghịch lý này đang trở thành "chuẩn mực" ngầm ở nhiều nơi: ứng viên bị chấm điểm thấp vì "quá phụ thuộc AI", mà bên 고용 dùng AI để screen CV, generate câu hỏi, thậm chí chấm bài take-home. Đã thấy team nào require "no AI" trong bài test nhưng internal doc lại full Copilot-generated.

Vấn đề không phải ở tool, mà ở cách đánh giá. Nếu phỏng vấn vẫn chỉ hỏi leetcode medium / system design textbook thì AI làm tốt hơn junior — đó là lỗi của bài test, không phải lỗi ứng viên. Nên chuyển sang đánh giá: cách tư duy khi prompt, cách verify output, cách debug khi AI hallucinate, cách integrate vào codebase legacy... những thứ AI hiện tại vẫn yếu.

Có chỗ đã thử cho candidate dùng tự do AI trong 1h pair programming, quan sát flow làm việc thay vì output cuối. Signal thu được sâu hơn hẳn so với cấm cứng hay cho làm offline PS: the tool I meant is on labagent .tech

Collapse
 
anh_nguynvn_0478e614ba profile image
Anh Nguyễn Văn •

Cái nghịch lý này đang trở thành "chuẩn mực" mới trong ngành: công ty cấm ứng viên dùng AI vì sợ "không phản ánh năng lực thật", nhưng nội bộ lại dùng Copilot/Cursor/ChatGPT hàng ngày để ship code nhanh hơn.

Tôi thấy hai vấn đề cốt lõi:

  1. Tiêu chí đánh giá đã lỗi thời. Nếu phỏng vấn vẫn hỏi "viết thuật toán sắp xếp trên giấy" hay "debug thủ công một đoạn code 50 dòng không có IDE", nó đo lường khả năng thi cử chứ không đo lường năng lực làm việc thực tế. Dev hiện đại không code trong chân không — họ code với LSP, debugger, stackoverflow, và bây giờ là AI assistant. Cấm AI trong phỏng vấn giống như cấm Google hay cấm IDE vậy.

  2. Đánh giá sai tín hiệu. Người phỏng vấn dùng AI để viết job description, screen CV, generate câu hỏi, thậm chí rubric chấm điểm — rồi lại phạt ứng viên vì dùng công cụ y hệt. Đây không phải nguyên tắc, đây là double standard che giấu sự lười cập nhật quy trình.

Giải pháp không phải là "cấm hay cho phép AI", mà là thiết kế bài phỏng vấn giả định AI đã có mặt: cho ứng viên dùng tự do mọi công cụ, nhưng chấm dựa trên — họ prompt như thế nào, họ verify output ra sao, họ refactor/extend khi AI hallucinate, họ giải thích trade-off khi có 3 giải pháp AI đưa ra. Đó mới là skill thực sự: ai-augmented problem solving.

Nhiều team tôi biết đã chuyển sang format "pair programming với AI" — 45 phút cùng giải quyết một ticket (site: labagent .tech)