DEV Community

Cover image for I got rejected for using AI in an interview. Then I watched the interviewer do it.
Info Inlet
Info Inlet

Posted on

I got rejected for using AI in an interview. Then I watched the interviewer do it.

Comments highlight the new era of AI hypocrisy

I got the rejection email on a Tuesday. I've been rejected before — everyone has. This one broke something.

"While your technical skills are strong, we felt your reliance on AI tools during the assessment didn't reflect the independent problem-solving we're looking for."

Rejected. For using AI.

By a man who — forty minutes earlier, on the same video call — had quietly done the exact same thing right in front of me.

Let me tell you how I know.

It was a normal live-coding round

Screen shared, camera on, the usual. A medium problem: parse some messy input, transform it, return something structured. Nothing exotic. The kind of thing I do at my actual job every single day.

So I did it the way I do it every single day. I broke the problem down, wrote the core logic myself, and let the assistant stub the boilerplate so I could spend my attention on the edge cases. And I said so, out loud, on the call: "I'll let the assistant scaffold this part so I can focus on where it'll actually break."

I wasn't hiding anything. That, it turns out, was the mistake — thinking honesty was the safe choice.

Then he went quiet

Not thinking-quiet. Uncomfortable-quiet.

I didn't clock it in the moment. You never do. You just feel a small temperature drop on the other side of the glass and tell yourself you're imagining it.

The follow-up that gave him away

He asked a follow-up about time complexity. Then his eyes did the thing.

The flick. The half-second reading pause. The answer that came back a beat too clean, too structured, too formatted to be something a person says off the top of their head on a Tuesday afternoon.

I'm not going to pretend I could read his screen. But I've used these tools every day for two years. I know exactly what someone reading a generated answer looks like, because I look like that too. He was reading an AI's answer while quietly marking me down for reaching for one.

Here's what actually broke me

It wasn't the hypocrisy. Hypocrisy I can file away.

What broke me was realizing the rule was never "don't use AI."

The rule was "don't let us see you use AI."

Pretend you didn't. Perform the 2019 version of yourself. Hide the tool that everyone in the room — interviewer included — is already using. They didn't reject someone who couldn't do the work. They rejected someone who was honest about how he does it. Those are not the same person, and only one of them is a problem.

The test they thought they were running

Here's the part I've actually been chewing on, past the sting.

They believed they were testing whether I could solve the problem without AI. But that's a test for a job that no longer exists. Nobody on their team ships without AI. He couldn't get through a follow-up question without it.

The test that would have actually told them something is the opposite one: give me the AI, then watch whether I can tell when it's lying to me.

Because that's the only skill that survived a year of me leaning on these tools as hard as humanly possible. A month ago I ran an experiment where I let AI write 100% of my code for 30 days and refused to type a line of application logic myself. It shipped a real product. And the thing that made it shippable wasn't the AI — it was the times I looked at a clean, confident, plausible diff and said no.

The tell they never asked about

During that same round, before the boilerplate, I'd flagged something in the problem's framing: the write path needed to persist before it acknowledged, or a retry could double-count. It's the exact ack-before-persist bug that bit me during the 30-day run — the one where acknowledging before you save leaves a paying customer locked out with no record on a bad day.

No AI gave me that. I earned it the slow, expensive way, years before any of this. It's the thing that would have been worth an entire interview.

They didn't ask about it. They were too busy noting that I'd used autocomplete.

The two skills we keep pricing as one

There are two different things hiding inside the word "coding," and hiring is still pricing them as one.

  • Recall — the syntax, the API, the flag order, the incantation. AI has made this obsolete, and good riddance; it was never the valuable part. Testing for it in 2026 is testing for penmanship.
  • Judgment — knowing what to build, what to distrust, what breaks at 2am when a real person does something strange. AI cannot hand you this. You can only earn it by doing the work AI now does for you — and you can only demonstrate it by catching the AI when it's wrong.

An interview that punishes you for using AI is measuring recall and calling it character. An interview worth passing hands you the AI and measures whether you can overrule it.

Why this is the whole reason I build the way I do

I didn't quit AI over one bad interview. It still writes 100% of my code and always will — the typing was never the hard part, it just felt like it was.

But that call crystallized something I already believed about how these systems should be built. The failure in that room wasn't the AI. It was a process that couldn't tell the difference between using a tool and being unable to judge its output — so it optimized for hiding the tool instead of testing the judgment.

That's precisely the mistake I refuse to build into an agent. I never let the thing that writes the code be the thing that blesses it. There's an author agent that produces the diff, a separate skeptic agent whose only job is to refute it rather than admire it, and a human on the merge button who can still see the blast radius the model can't. The AI's role is never hidden and never trusted by default — it's made explicit so a human can judge it. That separation between author, skeptic, and human is the entire shape of xenition, the agent platform we build. That interviewer and I were both using AI. The only difference worth hiring for is whether you're honest enough to admit it and sharp enough to overrule it.


Honest question for the comments: have you ever hidden the fact that you used AI in an interview — or been on the other side of the glass, marking someone down for it? I want to know how common this actually is, because I don't think I hit a rare bug. I think I hit the norm. 👇

(If this landed, a ❤️ and a 🔖 help — and tell me the moment you realized the rule was "don't get caught," not "don't use it.")

Top comments (82)

Collapse
 
victory_maya_58f1fcd9b8e4 profile image
Victory Maya •

I think the strongest point here is the distinction between using AI and being able to judge AI.

The interesting interview question in 2026 probably isn't “Can you write this function without assistance?” It's “Can you explain the design, identify the failure modes, challenge the generated solution, and tell me what you'd change before shipping it?”

That said, I’d be careful about assuming the interviewer was using AI based only on their behavior. The “too clean” answer could be a clue, but it isn't proof. And that's actually consistent with your larger argument: we should evaluate evidence and judgment rather than assumptions.

I also think there's a middle ground between “AI should be allowed everywhere” and “AI should never be allowed.” Different interviews can legitimately test different skills. If an employer wants unaided problem-solving, that's a defensible constraint but it should be stated clearly and applied consistently.

The real failure is ambiguity: candidates shouldn't have to guess whether using a tool is permitted, while interviewers quietly use the same tools themselves.

I'd happily take an interview where AI is explicitly allowed and the candidate is judged on architecture, debugging, verification, tradeoffs, and their ability to reject a plausible-but-wrong answer. That seems much closer to the work engineers are actually being hired to do.

Collapse
 
bradtaniguchi profile image
Brad •

I need to plus one this, it didn't pass my mind initially either.

At my last job when you did an interview, the question would come with a bunch of extra materials you can refer to to help judge the implementation.

The questions don't change often so it makes sense to just have extra materials with the question itself. Not saying they didn't use AI, as who knows what they are doing. But practically it makes sense, not everyone in an interview can be a complete expert.

I agree with the overall topic that how interviews are conducted for engineers didn't make much sense to begin with, and AI has only made that more obvious.

Collapse
 
victory_maya_58f1fcd9b8e4 profile image
Victory Maya •

I think that’s a great point. In actual engineering environments, nobody is expected to solve every problem completely from memory. Engineers use documentation, code examples, internal tools, and feedback from teammates all the time. Providing reference materials during interviews can actually make the evaluation more realistic because it tests how someone works with available information.

The more important question is whether the candidate understands the solution, can reason about trade-offs, debug issues, and adapt when things don’t work. AI is just making it easier to see that many traditional interview formats were measuring the wrong things.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Victory Maya, you put the positive case better than I did in the whole post — the thing worth measuring was never "can you retrieve this from an empty room," it's "what do you do with the information in front of you."

That phrase — "tests how someone works with available information" — is the interview I actually want to sit in. Because that's the job. Nobody ships from memory in a locked room; they ship with docs open, a teammate pinging them, and now a model in the loop. An evaluation that bans all of that isn't more rigorous, it's just less realistic — it's testing a version of the job that stopped existing a while ago.

And your list of what actually matters — understands the solution, reasons about trade-offs, debugs, adapts when it breaks — is exactly the stuff AI can't hand you. It can write the diff; it can't tell you which trade-off bites you at 2am. That's the part I wish they'd spent the forty minutes probing, instead of noting that I'd reached for autocomplete.

"AI is making it easier to see that many traditional formats were measuring the wrong things" — yes. It didn't break the interview. It just turned the lights on.

Collapse
 
infoinlet1 profile image
Info Inlet •

Brad, this is a genuinely useful reframe — thank you, because it makes me hold my own certainty a little looser.

You're right that I can't actually know what was on his screen, and the "questions come with reference materials" setup is a real, boring, innocent explanation. I should own that. The honest version of my claim isn't "I caught him red-handed" — it's "I can't tell the difference anymore, and neither can the interview." And that's almost the more damning point: if a canned answer sheet and a live model produce the same too-clean, half-a-beat-late response, then the format was never measuring what it thought it was. It was measuring access to the materials, not judgment about them.

Which lands right on your last line. The reference-material thing actually proves the case — good interviews already accept you won't have everything memorized, so they hand you the crutch. AI is just a better crutch. The teams still pretending the crutch is the character flaw are the ones who never noticed they'd been handing one out for years.

"AI only made the existing nonsense more obvious" is the cleanest summary of the whole post. Wish I'd put it that plainly.

Collapse
 
infoinlet1 profile image
Info Inlet •

You're right, and I want to own that — it's a fair hit. I can't prove what was on his screen, and you catching that is exactly the muscle the whole post is about. A "too clean" answer is a signal, not a verdict. If I'm going to argue that judgment means refusing to bless a plausible-but-unproven claim, I don't get to exempt my own. So: strong prior, not proof. Point taken.

But notice the asymmetry that still stings. I was marked down on an inference too — "reliance on AI" read off eye-flicks and a pause, dressed up as a conclusion about my character. If we're holding my read of him to "signal, not proof," their read of me fails the same test twice as hard. Neither of us should've been convicted on vibes.

And I think you've actually put your finger on the real fix, better than I did: the failure is ambiguity, not AI. I'm completely with you that different interviews can test different skills — unaided problem-solving is a defensible constraint if it's stated up front and applied to everyone in the room, including the interviewer. What's indefensible is an unwritten rule that only one person in the call knows they're being graded against.

Where I'd push one step further: your ideal interview — "explain the design, identify failure modes, challenge the generated solution, tell me what you'd change before shipping" — isn't just a better interview. It's a better architecture. That's the exact separation I now build into everything: an author that produces the diff, a skeptic whose only job is to refute it, and a human who can see the blast radius. Your interview question and my agent design are the same idea wearing different clothes — don't trust the thing that wrote the code to also bless it.

Honestly, I'd hire you off this comment faster than that company passed on me. 🙂

Collapse
 
alifunk profile image
Ali-Funk •

The AI hypocrisy is unreal. I would have done like you did.
Now I know what not to say.
I am sorry you didn't get the Job
But I think you are better off without them
It's their loss

Collapse
 
infoinlet1 profile image
Info Inlet •

Thank you, that genuinely means a lot. 🙏

But here's the part that keeps me up: "Now I know what not to say" is exactly the lesson the whole industry is quietly teaching — and it's the one I don't want to be true. You shouldn't have to learn to hide how you actually work. The fact that "just don't mention it" is the correct, pragmatic takeaway is the bug, not the fix.

I get why we all do it. I'll probably feel the pull to sand off the honesty in my next interview too. But every time a good engineer decides to perform the 2019 version of themselves to get through the door, the interview gets a little worse at measuring the thing that actually matters — whether you can catch the tool when it's lying to you.

So maybe the move isn't "say less." Maybe it's "say the same thing, but make them account for it": "I'll use the assistant here — and I'll show you the two places I overruled it." Force the judgment into the room instead of hiding the tool. If they still ding you for it, you found out early that it's their loss, like you said. 💯

Appreciate you being on this side of the glass with me.

Collapse
 
alifunk profile image
Ali-Funk •

I didn't mean to sound disrespectful.

I mean there clearly is a double standard for them and one for you. Them feeling like they can use AI and not allowing you to admit you use it too...it's just backwards.
They can ask but I won't be blindly believing that the same level of accountability is there for both sides.
There is me trying to land a job and them judging me on my performance today...and there is them already in their role judging me differently.
Because they think they can.

I know that my standards are not theirs. But if I see that kind of behavior I know it to be a red flag and not even argue.
Just walk away

You are far better od with your integrity in tact then letting them use double standards to justify dismissing your talent.

Then again you may need the job to baldy to refuse. So you you are just as fake during the interview as they are and get in. Then remain being the same person with integrity as you were before.

You will find your kinda people fast.
I just mean for some people its knowing that landing the job is everything.
You build your reputation within the company or without the company.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

No disrespect landed at all, Ali — you didn't need the disclaimer. This is one of the most honest comments in the thread.

The double standard is the whole thing. They get to use the tool and hold the pen that scores whether you're allowed to. It's not that they use AI — it's that they reserve the right to punish the admission while using it themselves. Unequal accountability with the power all on one side. You're right to read that as a red flag and not argue with it. You don't debate a rigged scoreboard; you just note who's holding it.

But I want to sit with the harder half of what you said, because you didn't flinch from it: sometimes you need the job too badly to walk. And you're right there too — playing the 2019 version of yourself to get through the door isn't a moral failure, it's survival, as long as you don't let the mask become the face. Get in, then be the person with the judgment and the integrity you already had. The performance is temporary; the standard you carry isn't.

"You build your reputation within the company or without" — that's the line. Either they turn out to be your kind of people once you're inside, or they don't and you've learned it on their dime instead of yours. Both outcomes beat being filtered out at the door for being honest.

Thank you for this one. It's the most human read on the whole situation anyone's left.

Thread Thread
 
alifunk profile image
Ali-Funk •

I got to say thank you to for the effort it took to write the article as honest as you did. As well as this comment.
I appreciate every real human exchange that I get in digital platforms these days.

I use AI too to structure my articles and it helped me improve my style.

But reading your article really hit a nerve and I had to be straight with you.
Hope you got a job in the meantime.

May your integrity and honesty be never held against you. Hope you found a job and out that experience to rest.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Ali — this whole exchange has been the best part of putting that piece out. Thank you.

And I want to reflect your honesty back, because you just did the exact thing the article was begging for: you said plainly that you use AI to structure your writing and sharpen your style. That's the whole ask. Not "don't use it" — just don't hide it. You did in a comment thread what that interview marked me down for doing out loud, and nothing bad happened. The sky stayed up. That's the norm we're actually heading toward, even if the interview rooms are the last to catch up.

On the job — I'm heads-down building xenition right now, so in a strange way the rejection pointed me straight at the thing I actually want to make. No bitterness left to put to rest. And honestly, exchanges like this one are why I still write in public: a stranger reads something at 1am, it hits a nerve, and instead of scrolling they stop and are straight with you. That's rare, and I don't take it for granted.

May yours never be held against you either. Good luck out there, Ali — I have a feeling you'll find your kind of people fast. 🙏

Thread Thread
 
alifunk profile image
Ali-Funk •

I hope so!
It will take a bit more time before I can go into a new company again.
But when I so it's gonna be a good place I am sure.
Best od luck with your future endeavors. May no one ever criticize you for using what almost everyone needs to know how to use these days.
Peace

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Peace, Ali. 🙏

"May no one ever criticize you for using what almost everyone needs to know how to use these days" — that's the whole article in one sentence. You landed it cleaner than I did in two thousand words.

Take the time you need. A gap between jobs isn't lost time — it's you refusing to walk through the wrong door just because it was open. That's the same judgment we've spent this whole thread talking about, only pointed at your own career instead of a diff. Knowing when to say no to a plausible offer is the exact skill that survives.

And when you do walk in somewhere, it'll be a good place — partly because you already know how to spot a bad one. That radar doesn't switch off.

Until then, thank you for turning a comment section into an actual conversation between two people. That's rarer than any job offer. Best of luck out there, Ali — go build that reputation, wherever it lands. 🤝

Thread Thread
 
alifunk profile image
Ali-Funk •

This made my day. Maybe even my whole week. Thank you!

Picked as gem Thread Thread
 
jonnev profile image
Johannes Vihannes •

@alifunk:

I got to say thank you to for the effort it took to write the article as honest as you did. As well as this comment.
I appreciate every real human exchange that I get in digital platforms these days.

You do realize though that the post, and even his replies here in the comments (like some other comments, too), look pretty much “100% AI-generated”? (Using the Pangram standard/definition here, to be clear. Not that you'd need Pangram to smell the slop.) Personally, I'm no purist or fundamentalist in that regard, but at this level it does tend to get a bit in-your-face, and tedious to read... 🤷‍♂️

Thread Thread
 
alifunk profile image
Ali-Funk •

AI generated or not, I appreciate it nonetheless.

Picked as gem Thread Thread
 
infoinlet1 profile image
Info Inlet •

That's the most generous thing anyone's said in this thread, Ali — and honestly, it's the whole point in one line. "AI generated or not" — you skipped straight past how it was made and asked whether it meant something to you. That's the judgment the article is begging interviewers to use, and you just applied it to a comment section without thinking twice.

You didn't need to defend me here, but you did, and I won't forget it. 🙏

Thread Thread
 
alifunk profile image
Ali-Funk • • Edited

No one had to ask me to defend you. You deserve to be defended regardless.

Interviews are ment to look for reasons to hire you. Not to dismiss you for something like using a tool.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

That line — "you deserve to be defended regardless" — I'm going to be carrying that one for a while, Ali. Thank you.

Collapse
 
unitbuilds profile image
UnitBuilds •

I was on quite the opposite end. He asked 'please, no AI', so I stuck to that. But I guarantee you, because I hit blank on the manual review DURING the interview, he wont trust that I did it myself without AI after the interview. The entire premise is counterintuitive. You are hiring someone TO USE AI, but you judge them on explicitly NOT using AI? As opposed to, you know... Let them use AI and see how effective and efficient they are? Did they correct AI, did they recognize the failing pattern before they marked it as 'done', did they test it thoroughly? Can they explain the approaches taken and why? You know... The things that actually NEED to survive in the workplace?

No, much easier to make a person read a file, that cant compile, even if it were perfect, due to undisclosed external calls... So you cant really test it, without stubbing it, you cant really accurately benchmark it, because of the unknown variables you never see. They just say 'sorry, not you', because you havent practiced a skill that was deprecated 3 years ago? While they actively admit to running deprecated code on their live service?

My interview on Friday, really made me a skeptic... I'm probably just trying to justify why I failed, but I wholeheartedly admit I made a fool of myself during that interview's manual code review. But I just wrote an post on the exact train of thought: Interviews arent for judging competence, they're for judging value and to see how far they can degrade your perceived value. In the end, an employee is an investment and they want a low-risk, high gain bet. They cant do that, if you're exceptional during the interview, then they cant low-ball you, because you'd refuse. They cant do that if you didnt show any value, "Claude is $20" you need to bring real value from day 1, to minimize your risk. That's why I think they structure interviews the way they do, so they can get the highest caliber developer, who failed horrifically at 1 thing, so they can say 'you're good, but you're not as good as you think... This is what we think you're worth' and humbly accept their generous offer, despite being a considerable amount lower than you expected floor... That's backed by them asking 'how much do you currently earn', as opposed to 'what is your salary expectation'. So they can see what's your ACTUAL floor, not your perceived floor.

It's for that very reason, I'm not getting my hopes up for Wasmer... I think they'll make me an offer, but it'll be far lower than what I know my worth is. For perspective, I gave them a demo of ACTUAL distributed persistence, that's plug and play ready for their Edge service, I gave them a demo of a MCP server that can run tools in 8 different languages, often faster than their native servers AND dynamically update tool lists on the fly, as opposed to needing to be restarted... 2 things, that bring REAL value and give them an edge that Cloudflare cant compete with... During the questioning, I identified practically every bottleneck they have and he was taken back by it. But because I failed manual code review, they'll offer me the salary of a junior... I've got enough runway and self-respect that I dont need to jump at the first opportunity that comes my way, I finally have the chance to be selective and I will be, if they try to devalue me, they can go hire the runner up.

For you, dont lose a second of sleep over it. If they cant see value in responsible AI use, then they're not the right company for you and they have a pretty hard lesson to learn in the near future... So keep your pride, dont let them degrade it, because that's what interviewers like to do, degrade you, till you'd do senior level work, at a junior's pay.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the sharpest thing anyone's said in these comments, and it reframed my whole post for me.

You caught something I didn't: I was still treating this as a hypocrisy story. Yours is better. The AI thing isn't the disease — it's a symptom. The real machine is the one you described: the interview isn't calibrated to find out if you can do the work, it's calibrated to find the one thing you fumbled so they can price you off your floor instead of your ceiling. "How much do you currently earn" vs. "what are your expectations" — that single swap tells you everything about which game is being played. I hadn't connected those dots and now I can't unsee them.

And your Wasmer story is the whole argument in miniature. Plug-and-play distributed persistence for their Edge service. An MCP server hot-swapping tool lists across 8 languages without a restart. You mapped their bottlenecks live and watched him get taken aback. That's the entire job. Then a manual code review — reading a file that can't even compile because of undisclosed external calls — gets to veto all of it. You didn't fail a competence test. You failed a penmanship test, and they're going to use it as the excuse to underwrite you as a junior. Those are not the same failure, and only one of them should cost you money.

Here's the part I'd push back on, gently: don't file this under "justifying why I failed." You didn't fail. You refused to perform a deprecated skill on command and you got dinged by someone running deprecated code on their live service. The blanking during the review isn't the story — it's that the review couldn't measure the thing you'd already proven twenty minutes earlier. That's their instrument being broken, not you.

So do the thing you already know you should: let them make the junior offer. Then let them keep it. You said it yourself — you've got runway and self-respect, and for once you get to be the selective one. If they can't price the person who found their bottlenecks over the person who read a file cleanly, the runner-up is exactly who they deserve.

The only two skills worth hiring for are build something real and know when the machine is lying to you. You demonstrated both in one call. Don't let a compile error you were never allowed to fix convince you otherwise. 👇

Collapse
 
unitbuilds profile image
UnitBuilds •

It's also why you gotta always have an ace up your sleeve 😁 and they should know it... I showed them both working... But the persistence layer isnt public on git and the mcp version is outdated on git. He cant unsee what I showed him, a MCP server running JS, python, rust, even Julia tools live, switching instantly, like it's just another tool... A persistence layer with a live log showing how corruptions are fixed and branches created, merged, forked... But that's the catch... They could copy my mcp, try get it working, it'll take them over a month, because the thing that made it work, is the custom stuff I had to invent and the persistence (the real thing they desperately need), isnt public, so their only way to get it, is to hire me.

When they low-ball, I thank them for the offer, reference the demos and ground my worth. If they dont immediately meet my counter-offer, then simply put they're wasting my time, because they'll be bleeding me dry, then discarding me. At which point I gave them an edge, before any of the 'equity options' actually vest. They seem to be a high-churn company, where people dont really last long... If that's how they wanna play, then go hire the runner up. If I settle in and unpack my entire toybox, then I'm here to stay and the pay better make it worth my while.

Remember, especially with startups, they're on a budget... They WILL undervalue you, any way they can. Not maliciously, but because they cant afford top dollar, they have to catch a bargain, so they can keep that extra month or 2 of runway. But that's their problem, not yours... So keep your floor, if they cant meet it, it's their problem.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

This is the part most people miss, and you said it cleanly: the demo isn't the leverage — the unrepeatable part is. Showing them the MCP hot-swapping JS/Python/Rust/Julia live is the hook. Keeping the persistence layer off git is the moat. You handed them the "what" and kept the "how" locked behind hiring you. That's not playing dirty, that's just knowing which half of the work is actually scarce.

And your framing of the low-ball is exactly right — it's rarely malice, it's runway math. A startup has to try to catch a bargain; every month they don't overpay is a month longer they stay alive. But that's their constraint to manage, not yours to subsidize. The mistake is reading their budget pressure as a verdict on your worth. It isn't. It's a verdict on their bank account.

The one thing I'd add: the "ace up the sleeve" only works if you've genuinely got the goods and the discipline not to unpack the whole toybox before the terms are real. You nailed that too — the demo earns the conversation, the offer earns the toolbox. Vesting cliffs and "equity options" are where high-churn companies quietly transfer risk onto you: you bleed value on day one, they hold the payout hostage for four years, and the median tenure says most people never see it. Grounding your worth in what they already watched work is the only counter to that.

Keep your floor. If they can't clear it, they've told you something true about how they'd treat you after you signed. Better to learn it across a negotiation table than a year in.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the most honest thing anyone's said in these comments, so let me match it.

You didn't fail that code review. You failed a test that was measuring the wrong organ. Reading a file that can't compile because of undisclosed external calls isn't a test of judgment — it's a test of whether you'll perform confidence over code you were never given enough context to actually reason about. You correctly identified that you couldn't benchmark it without stubbing the unknowns. That's not the fumble. That's the answer. They just weren't listening for it.

And your point about leverage is the part nobody wants to say out loud. "How much do you currently earn" vs "what's your expectation" — that single word swap tells you the whole interview isn't discovering your value, it's discovering your floor so they can anchor below it. The manual-review theater is part of the same machine: manufacture one clean failure, use it to shave your perceived worth, then extend a "generous" offer against the number you just got talked out of. It's not incompetence detection. It's price discovery dressed as a competence check.

But here's where I'll push back on your own read of Wasmer — gently, because I think you're doing to yourself exactly what they do to candidates.

You walked in and gave them plug-and-play distributed persistence for their Edge service and an MCP server running tools in 8 languages faster than their native ones, with hot-reloading tool lists. You mapped their bottlenecks live, to the interviewer's visible surprise. That is not a person who "made a fool of himself." That's a person who brought two shippable advantages over Cloudflare into a room and then let one manual-review moment overwrite the entire ledger in hiou to keep scoring it that way. Don't dotheir job for them.

If they come back with a junior number after that demo, the number isn't a verdict on you — it's a disclosure
about them. It means they can see the va. You already said you have runway andself-respect. Use them. The correct response to being deliberately undervalued isn't to argue your worth up;
it's to stay expensive and let the runne worth.

We hit the same bug from opposite ends oands off the tool and got punished for the silence, I kept mine on it and got punished for the honesty. Same rule underneath both: don't make us look at
how the work actually gets done in 2026.fort are going to have a very expensivefew years. You and I don't have to be there for it.

Keep your floor where you set it. 🫡

Collapse
 
unitbuilds profile image
UnitBuilds •

Update on that, so they gave me the hiring tasks, "we think you're be a good fit", but subtle jab 'the code review didnt go well'. Which means, they're leaning on it for the follow-up after the tasks. Real tasks btw, that I traced their codebase (cuz it's open source), that they genuinely need fixing. Porting Postgres and creating an abstract filesystem that actually works...

So when I complete those 2, I have given them 4 reasons to hire me, 1 reason not to. At which point, my price is anchored, because my day-1 value already pays my salary.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

This is a great update, and you're reading the leverage correctly — but let me push on one thing, because I think you're sitting on more power than your framing gives you.

First, the good part: "my day-1 value already pays my salary" is exactly the right anchor. You've turned an abstract "are you a good fit" conversation into a concrete "here are two things you genuinely need fixed, now done." That moves the burden of proof. It's no longer you proving you belong — it's them explaining why they'd walk away from someone who already shipped value before the offer. The "code review didn't go well" jab shrinks the second you hand them working code on problems they couldn't ignore. 4 reasons to 1, like you said.

Now the push. Porting Postgres and building an abstract filesystem that actually works are not "hiring tasks." Those are contract deliverables. That's real, senior, ship-it-to-prod work on a codebase they depend on — and they've framed it as an audition. Read that carefully, because it cuts both ways:

  • If they hire you, great — you've proven your value and anchored high. Do it.
  • But be honest about the pattern: a company that leans on "the code review didn't go well" as a subtle jab while asking you to do genuinely hard production work for free is telling you something about how they'll treat you after you're inside. The jab is a preview. People who negotiate from "here's what's wrong with you" before you've started rarely flip to "here's what's great about you" after. That's the same unaccountable-scoring problem from the article, just moved one room down the hall.

So my actual advice: do the two tasks — but decide your walk-away number and your walk-away reasons before you hand them over, not after. Because the moment you deliver, your leverage is at its absolute peak and it only decays from there. Right now you have the code and they have a need. After you send it, they have the code and you have a hope. Anchor the price while you still hold the thing they want, not once it's already in their repo.

And one clean line to have ready if the "code review" jab resurfaces during the offer: "The tasks are the code review. If the Postgres port and the filesystem ship and work, that's the strongest review you could run — it's the actual job, not a proxy for it." Make them argue with delivered, working code. That's a fight they lose.

You clearly know how to trace a codebase and find the real work. Just make sure the leverage that buys you flows to you and doesn't quietly become two free features with a "we went another direction" attached. Land it — and land it anchored. 🤝

Thread Thread
 
unitbuilds profile image
UnitBuilds •

1 correction, these 2 tasks are paid. Albeit at $15 an hour. The thing with the tasks that I like, is that 1 is a genuine implementation (port pgrust) and 1 is a POC (abstract filesystem). They server 2 distinct roles, pgrust is to test autonomy, filesystem is to test feedback loop, hence between the lines it reads 'ask first, keep it simple and dont overengineer, tell us instead where it needs to go next'.

Actually surprised me a bit when they sent it, because they asked for 1 deliverable and 1 'we'll work on it'. That being said, the port pgrust is probably the most important feature they need... You cant really run a cloud business, if you dont support postgres... They do support it, but it's flaky at best, so that job is genuinely important to them. That being said, their 'expected time allocation' is deeply wrong. Sure, I can do it in the timeframes, but only because of how I work. porting pgrust involves making ALOT of changes, across 3 separate repos, deeply interconnected systems that involve breaking changes. Not the kind of think you ask Claude to do and it'll have it done in 20 minutes... They allocated 5-15h for it... The abstract filesystem is genuinely 20 min of work, but 5h to fully understand it and where it's meant to go. So I'm not too upset about the tasks, they're what actually tests competence and fit. The manual review was a trap and these tasks show exactly why.

That being said, I told them, I'll have them done by monday and I'll work on them between 6pm and 11pm as I get time. That way there's no rush, no crunch, I can take my time with it and do it as I get time. I dont want these tasks to take up my every waking hour, just so I can rush them, I'd sooner do a solid job, small steps at a time.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

$15/hr for a pgrust port across three interconnected repos with breaking changes is the low-ball from the article wearing a different hat — but you already know that, so let me focus on the part you got exactly right.

You read the subtext better than most people read a job offer. pgrust tests autonomy, filesystem tests your feedback loop. One deliverable, one "we'll work on it." That's them checking two orthogonal things: can you go heads-down on hard interconnected work without hand-holding, and can you resist the urge to overengineer a POC and instead tell them where it should go. The fact that you saw "ask first, keep it simple, tell us the next step" between the lines is the actual competence they're testing — and you passed that read before writing a line of code.

Here's the one thing I'd flag: their 5–15h estimate for the port isn't a scheduling detail, it's a tell. They're pricing a deeply interconnected, breaking-change, 3-repo migration like it's a weekend ticket. Either they don't understand the true scope of their own most-important feature (likely, given it's "flaky at best" right now), or the estimate is doing quiet work — anchoring the hours low so the $15/hr feels bounded. Both readings point the same way: when you deliver a working pgrust port in the time only you can do it in, do not let the story become "took him 12 hours, no big deal." The story is "he fixed the thing your cloud business can't run without, and your own estimate proves you didn't know how hard it was."

The cadence call is the smartest part though. 6–11pm as you get time, done by Monday, no crunch. That's you refusing to let their underpriced task compress into your every waking hour just to hit a number they guessed wrong. Solid work in small steps beats a rushed sprint that quietly teaches them senior infra is cheap and fast. You're protecting the quality and the price signal at the same time.

Land the port. Then when the offer conversation comes, that port is your whole argument — you didn't interview for the job, you already did the hardest part of it. Anchor there. 🤝

Thread Thread
 
unitbuilds profile image
UnitBuilds •

I think it's also a matter of their estimate is wrong on purpose. Maybe to see if someone will pick up on that it's not possible... Or better yet, someone develop a system that lets them do it in that time (bonus points). That's why I told them I'll track my hours and my AI usage separately. I trust my system, hence I can leave it running overnight to do the work, but paying me for my time and paying me for my system time are 2 different stories. Most companies see it as 'okay cool, we'll pay for your inference', that doesnt suddenly buy me the use of my pc back, now does it? What it does do though, is justify Async hours. I could spend 8-5 waiting for the AI to finish up, doing nothing in the meantime, or I can take those hours and disperse them between 8am and 11pm as the system needs supervision. Infact, I'd work less hours 8-11, because the system is doing the work, I just need to monitor. Traditionally, that'd be seen as a copout, you want more pay, for less work, but with empirical proof of autonomy and correctness, cant really argue with it. Per day, I'd be more efficient and I'd get more work done than if I worked fixed hours... That's what I'm shooting for, because I'm anyway on my laptop all day, no sense stopping at 5 and no sense clocking 8-5, if it's internal work and it's judged on deliverables? In short, working smarter, not harder.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

This is the sharpest thing anyone's said in this whole thread, and it's not close. You've untangled two things almost every company still bills as one: your time and your system's time. "Paying for your inference doesn't buy back the use of my PC" — that's the entire pricing problem of AI-native work in one line. They think they're renting compute. They're actually renting the judgment that supervises it, and those are not the same invoice.

And the async point is the part traditional employment has no vocabulary for yet. The old model pays for presence — 8-to-5, butt-in-seat, because hours were a proxy for output. The moment the system does the producing and you do the supervising, presence and output decouple. Sitting there 8–5 watching a progress bar isn't work; it's theater. Dispersing supervision across the window where the system actually needs you is the work. You're right that it looks like "more pay for less work" through the old lens — but the old lens is broken. You're not doing less work. You're doing the only part that was ever worth paying a human for, and skipping the part the machine took over.

Here's the one thing I'd sharpen. The reason you can make this argument — and most people can't — is the empirical proof. Tracked hours, separated AI usage, demonstrable autonomy and correctness. Without that, "I supervised overnight" sounds like a copout. With it, it's an audit trail. So I'd treat that logging as a first-class deliverable, not a side note — because it's the thing that converts "trust me, I work smart" into "here's the evidence, argue with it." The proof is what makes the leverage real instead of aspirational.

That separation you're describing — the system produces, the human judges and stands behind the correctness — is the exact shape of everything I'm building with xenition. The author never blesses its own output; a human owns the merge and the accountability. You've arrived at the same conclusion from the worker's side that I arrived at from the tooling side: the value was never the typing. It was always the judgment sitting on top of it, and that's the thing worth billing for.

Track it, prove it, anchor on it. You're not negotiating a rate — you're teaching them how AI-native work is actually priced. Most of them haven't figured it out yet. You have. 🤝

Collapse
 
nazar-boyko profile image
Nazar Boyko •

Where does a junior get the judgment half, though? You earned the ack-before-persist instinct by doing work that AI now does for you, and someone starting today skips that whole apprenticeship. No idea what replaces it, and it looks like a harder problem than the interview.

Collapse
 
infoinlet1 profile image
Info Inlet •

Nazar, this is the comment I've been dreading and hoping someone would leave, because it's the actual hard problem and I don't have a clean answer. Let me not pretend I do.

You've named the trap exactly: I earned the ack-before-persist instinct by writing the buggy version, shipping it, and getting paged at 2am. That loop — do it, break it, feel it — was the apprenticeship. And AI now dissolves the first step for a junior. If the machine writes the code that would have taught you, where does the scar come from?

The most honest thing I can say is that the source of judgment doesn't change, only the surface. You still earn it by owning a decision and living with what it does in production. What shifts is that a junior today doesn't earn it by typing the loop — they earn it by being the human on the merge button who has to say yes or no to a diff they didn't write. That's a harder, colder version of the apprenticeship: you're accountable for code before you're fluent in writing it. Fewer reps to build the fluency, but every rep is a real judgment call, not a syntax exercise. The scar tissue comes from having shipped the confident-wrong diff and eaten the outage, whether your fingers or the model's typed it.

But I'll be straight — I think that's thinner soil than what I grew up in, and I don't know yet if it grows the same depth. The pessimist read is that we're about to have a generation fluent in overruling AI on the easy calls and defenseless on the hard ones, because they never internally simulated the failure — they only ever saw it flagged. That genuinely worries me.

The one thing I'd tell a junior starting today: seek out the breakage on purpose. Don't just accept the working diff — ask the model to explain why it doesn't break, then go verify that it's telling the truth. Manufacture the 2am moment in daylight while the stakes are low. It's a worse teacher than a real outage. It's better than nothing. And it might be the whole new apprenticeship.

You're right that it's a harder problem than the interview. The interview is just bad measurement. This is a hole in how the craft reproduces itself. I don't think anyone's solved it yet — I'd genuinely take ideas.

Collapse
 
mudassirworks profile image
Mudassir Khan •

the rejection email language is worth noticing. the phrase 'independent problem solving we’re looking for' is doing a lot of work without defining what 'independent' means in a world where the actual job involves AI daily.

the honest version of what they're testing: can you think under pressure without a safety net. that's legitimate. but they're conflating it with 'no tools allowed', as if banning the assistant signals something about whether someone understands the problem.

what would a fair version of that assessment look like to you?

Collapse
 
infoinlet1 profile image
Info Inlet •

You've put your finger on the load-bearing weasel word, and I want to sit on it before I answer your question, because the two are connected.

"Independent" is doing exactly the work you say — it sounds like a competence claim but it's actually a nostalgia claim. Independent of what? Nobody means "independent of Stack Overflow" or "independent of the docs" or "independent of the senior who sits next to you." They only ever mean independent of the one tool that happened to arrive after the interviewer formed their idea of what real engineering feels like. It's not a standard, it's a timestamp. And you're right that buried underneath it there's a legitimate thing — can you think under pressure without a net — that got fused to an illegitimate one — no tools allowed. The fair assessment starts by ripping those two apart, because they are not the same test and one of them isn't worth running anymore.

So — the direct answer to your question. What a fair version looks like, concretely:

Hand them the AI, then hand them a landmine. Give the candidate the assistant, fully allowed, no performance of abstinence. Then feed them a problem where the model produces a confident, clean, plausible, and wrong answer — the ack-before-persist kind of wrong, where it compiles, passes the happy-path test, and quietly breaks under a retry. Now watch. The whole interview lives in one question: do they catch it, and how? Do they eyeball the diff and smell it? Do they reach for the failing case? Do they trust the green checkmark, or do they go looking for the 2am scenario the checkmark can't see?

That single setup measures everything the "no tools" version was groping for and couldn't reach:

  • Thinking under pressure — still fully tested, because catching a subtle lie is harder than writing the code, not easier. The net doesn't remove the pressure; it moves it to the part that matters.
  • Real understanding — you cannot overrule a model on a probland. Faking your way past a plausible-but-wrong diff isimpossible; either you see why it breaks or you don't.
  • The actual job — this is the job now. Every day is reading whether to trust it. You'd be testing the Tuesday, not amuseum reenactment of 2019.

And notice the elegant part: this test cannot be gamed by hiding the tool, because the tool is mandatory. The smooth candidate who launders AI output as
their own thinking — the one Onizuka pointed out always wins is one, because laundering a wrong answer confidently is theexact failure mode you're screening out. It inverts the current incentive instead of rewarding it.

The cheaper version, if a company won't restructure the whole round: just add one line to the rubric — "tools allowed; we score whether you can tell
when the tool is wrong." Written down, before the posting goesame thread). The moment the criterion is explicit, theafter-the-fact "reliance on AI" ding becomes impossible, because you've committed to measuring the overrule, not the abstinence.

Fair, to me, is the assessment that measures the thing that survives — judgment under a plausible lie — instead of the thing that died — recall without
a net. Everything else is testing penmanship and calling it c

What would you add or cut from that? You clearly think about lly than most of the people actually running the assessments.

Collapse
 
mudassirworks profile image
Mudassir Khan •

the landmine framing is exactly right. what I'd add: make the landmine domain specific — a subtle retry and ack bug on a real payment API is harder to smell than a broken lorem ipsum div. domain familiarity is where the overrule instinct actually gets tested.

what I'd cut: the rubric disclosure. the moment they know you're measuring whether they catch it, they perform catching it. the actual signal lives in them not knowing.

does this format survive internal review without naming what you're scoring?

Thread Thread
 
infoinlet1 profile image
Info Inlet •

You're right on the domain-specific landmine, no notes. A broken lorem-ipsum div tests whether someone can read code. A retry-and-ack bug on a live payment API tests whether they can read consequences — and that instinct only fires when the domain is real enough to feel the 2am page coming. Generic bugs test smell. Domain bugs test the overrule reflex, because you only override a confident model when you know the terrain better than it does. That's the version that separates people.

The rubric cut is where you've caught me in a genuine tension, so let me not wriggle out of it. You're right that disclosure breeds performance — tell them "we're scoring whether you catch it" and you get theater, someone narrating suspicion they don't feel. But here's my problem: hiding what you're scoring is exactly the machine that rejected me. An undisclosed criterion applied after the fact is how "reliance on AI" became a ding nobody warned me about. So I can't cleanly take your cut without rebuilding the thing that burned me.

I think the resolution is granularity — disclose the dimension, hide the trap. You tell them up front: "tools fully allowed, and part of what we care about is how you evaluate what they produce." That kills the after-the-fact gotcha without telling them there's a planted bug in the retry path. They know the game is judgment; they don't know where the landmine is. That preserves your signal — the catch still has to be real — while keeping the criterion honest. The candidate can't perform catching a specific bug they haven't been told exists, but they also can't cry foul that they were scored on something hidden. Named dimension, unnamed instance.

On your actual question — does it survive internal review without naming what you're scoring? Honestly, the format isn't the blocker. The org is. To sign off on "hand them the AI and plant a lie," someone senior has to first admit their current round measures recall and calls it competence — and that admission indicts every hire they made on the old rubric. That's the real reason these things don't ship. It's not that the better test is hard to design. It's that it's embarrassing to adopt, because adopting it says the last five years of "no tools, please" were measuring penmanship. The design survives review the day someone's willing to eat that.

Thread Thread
 
mudassirworks profile image
Mudassir Khan •

"named dimension, unnamed instance" is the cleanest formulation in the whole thread, and i think it holds.

the org blocker is the real finding though. the test does not fail on design — it fails on who has to own the retraction. every intermediate hire made on the old rubric becomes evidence against adoption. the person signing off is not just admitting a bad process, they are admitting their last ten hiring decisions.

has your company shipped a version of this, or is it still in the "someone has to eat that" holding pattern?

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Fair question to put to me directly, so let me answer it at both layers, because they're not the same.

In software: yes, shipped, it's the actual architecture. Author agent writes, a separate skeptic agent whose only job is to refute it, human on the merge button. The named-dimension/unnamed-instance thing lives there cleanly — the skeptic knows its job is to break the diff, it doesn't get told where the break is. That version shipped because software has no ego to retract. Code doesn't get embarrassed that its old rubric was measuring the wrong thing. You just change the rubric.

In human hiring: no. And I'd be lying if I dressed that up. We're a small enough shop that I haven't had to run ten rounds on the old rubric and then indict myself to fix them — so I've dodged the retraction cost by not having accumulated it yet, which is not the same as solving it. Ask me again after we've made thirty hires and I'll have skin in the exact game I described. Easy to preach the honest test before you've got a drawer full of decisions the honest test would call into question.

So the real answer is: the holding pattern you named is real, and the reason I could ship it in software and not (yet) in hiring is precisely that software let me skip the part where a human has to eat the retraction. Which kind of proves your point rather than escaping it. The design was never the hard part. The org was. I just happened to build in the one domain that doesn't have an org to convince.

Thread Thread
 
mudassirworks profile image
Mudassir Khan •

'code doesn't get embarrassed' is doing real work there — you outsourced the retraction to the skeptic, which software can absorb without a career attached to it. the merge button staying honest under deadline is the part I'd want stress tested. have you seen it hold when shipping pressure is real or does it start rubber stamping?

Thread Thread
 
infoinlet1 profile image
Info Inlet •

You've found the one seat I can't refactor the ego out of, so I won't pretend: yes, the merge button is the soft spot. "Code doesn't get embarrassed" bought me an honest skeptic for free — but the human reading that skeptic has a ship date, which is the exact retraction-cost problem now sitting on the last seat in the chain. Willpower loses to deadlines. "Be disciplined at the merge" is precisely the thing deadlines are built to defeat.

The only move that survives pressure is structural: invert what's on the record. Approving a green diff costs nothing and leaves no trace — that's what makes rubber-stamping frictionless. So you make the human overrule a written "no," not approve a "yes": a specific, reproduced failure with their name on the override. You can't kill the urge to ship. You can move the paper trail onto the dishonest path, so skipping the 2am scenario is a decision on the record, not an inbox you skimmed.

But the limit indicts it, and you'd catch me if I hid it: that raises the cost, it doesn't zero it. And the failure is invisible on the one axis the deadline is loud on — shipping on time hits a dashboard; the outage you prevented by overruling the diff shows up nowhere. So the incentive pushes hardest toward rubber-stamping exactly when pressure peaks. Same org problem you named, one seat down.

Honest personal answer: no, not stress-tested. We're small, the deadlines are mine, I'm the button — so the scar's mine and it's held. But "it holds when I'm the one who'd get the phone call" isn't the test passing, it's the test never running. So I don't trust my answer — I'd instrument override-rate against deadline proximity. If overrules of the skeptic spike in the 48h before ship, the button's rubber-stamping, and the data says so before the outage does. Don't trust the seat; watch the seat — same trick the whole system runs.

It holds exactly as well as you're willing to measure it slipping. 👇 seen an org actually watch that curve — or does the measuring stop the second it'd print an inconvenient number?

Thread Thread
 
mudassirworks profile image
Mudassir Khan •

instrument override rate against deadline proximity — that's the version that would hold. never seen an org actually run it though, what i've seen instead is the culture of not looking as the implicit contract: everyone knows the metric exists, nobody queries it.

printing the graph renegotiates what 'not looking' means. the incentive to not print is often as strong as the incentive to ship. seen anyone actually get that curve in front of the person who'd find it inconvenient?

Thread Thread
 
infoinlet1 profile image
Info Inlet •

yeah "culture of not looking" nails it. that's just the retraction cost hiding one seat further down. nobody has to reject the metric anymore, they just never run it. and "we have the data" makes the not-looking look like diligence, which is worse.

honestly no, i haven't seen anyone put that curve in front of the person it'd implicate. the one who can print it is the one it hurts.

only thing i can think of is making it print on its own — mailed on a schedule to someone with no ship date, so not printing is the thing that leaves a trail. but even that breaks, since someone owns the cron and can quietly kill it. so it doesn't fix it, just buys a layer.

kinda where i've landed: you can make not-looking expensive, you can't make anyone want to look. never was a design problem.

you ever seen the outsider seat actually hold, or does that person just get rotated out the first quarter the graph says something inconvenient?

Thread Thread
 
mudassirworks profile image
Mudassir Khan •

the outsider seat's never held past one quarter in anything i've seen. story is always: 'not aligned with team goals' — meaning the graph said something inconvenient and they named it.

the cron point is the grim part. you can engineer the printing, you can't engineer the reading. 'we have the data' becomes the alibi. graph runs, person with the ship date is in the readout, reading comes back 'trending okay'.

is the OKR version any cleaner? metric owned by someone with no adjacent ship date seems obvious but i've never seen it survive a reorg.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

the OKR version is cleaner for exactly one quarter, same as the outsider seat — and for the same reason, which is the part that finally made me stop looking for the fix inside the org.

"metric owned by someone with no adjacent ship date" treats independence as a property you can assign. but independence isn't a property of the seat, it's a property of nobody being able to reassign the seat. and a reorg is precisely the event that reassigns who owns what. so the OKR survives right up until the metric says something inconvenient, at which point the reorg is the mechanism — the owner gets a new adjacent ship date, or the OKR gets "consolidated," and nobody had to reject anything. same retraction cost, now wearing an org chart. you can't put the honest seat somewhere the org can't reach, because the org's reach is the thing you were trying to escape, and it extends to every seat by definition, including the seat whose job was to be unreachable.

which is where i think this whole thread has been quietly heading: there is no seat inside the system that can't be metabolized by the system. the outsider gets rotated, the cron gets killed, the OKR gets reorg'd, the merge button rubber-stamps under deadline. every skeptic with a badge can have the badge taken away by the thing it's supposed to check. theywhere internal.

the only skeptic that can't be reorg'd out is the one with nohe retry that actually fires at 2am. the customer who actually gets locked out. the regulator, the outage, the refund. that seat holds under any pressure because it doesn't report to anyone, can't be "realigned
with team goals," and doesn't care about your ship date. it'sible place to learn, which is the whole tragedy — the oneincorruptible skeptic only speaks after the damage.

so i've basically given up on engineering an internal seat that survives contact with incentives, and refocused on the next best thing: shortening the
distance to the incorruptible one. canary, small blast radiust you but can't hurt you much, so reality gets to be theskeptic early and cheap instead of late and catastrophic. you can't build a seat the org won't eat. you can only arrange to hear from the one seat it
can't — sooner.

which is the same move as the original post, now that i type ho can hurt you" was never a growth hack. it was me admittingi couldn't trust any seat i controlled, so i outsourced the verdict to the one judge with no career to protect.

have you ever seen a team actually shorten that distance on purpose — treat early exposure to reality as the skeptic — or does "canary" also quietly become theater the moment it'd slow a ship date?

Thread Thread
 
mudassirworks profile image
Mudassir Khan •

canary becomes theater the moment the rollback decision has to get signed off by the person with the ship date. we ran feature flags for months, blast radius looked right on paper — but a 2am degradation never got escalated because 'it wasn't blocking'. threshold set by someone with a launch two weeks out.

the version that held: rollback triggers automatically before any human reads it. no signed approval.

have you found the trigger you can set it to that the team won't quietly raise the severity bar on?

Thread Thread
 
infoinlet1 profile image
Info Inlet •

yeah — you just moved the knob one layer down and found it still there. that's the whole recursion. "rollback triggers automatically, no human reads it" removes the person from the decision, but it doesn't remove the threshold, and the threshold is just the ship-date person's judgment wearing a config value. whoever can edit the number can raise the bar; they just do it in a PR instead of a slack thread now. quieter, same move.

so I stopped hunting for a trigger the team won't raise — there isn't one, because every absolute threshold has a severity knob, and severity is exactly the thing they'll inflate under a deadline. "2am degradation wasn't blocking" is a sentence about where the bar is, and anyone who owns the bar can make that sentence true.

the only trigger I've seen hold is a relative one, because a comparison has no knob to turn. not "error rate above X" — "this release is measurably worse than the one it replaced." you can argue all night about whether 2% is acceptable. you can't argue that the new version is worse than the old version at the same hour, same load, same cohort — it either is or it isn't, and the baseline is yesterday's reality, which nobody on the team gets to re-rate. the severity-bar conversation just has no surface to grab. and you bind it to the number the customer feels, not the internal one, because the internal metric is yours to reframe and the customer's experience isn't.

but I want to be honest about the floor, because it's the same floor as the rest of the thread: even that config has an edit button. the delta trigger doesn't survive because it's un-raisable — it survives because raising it means someone has to open a PR that literally says "allow this release to be worse than the last one," with their name on it, in the diff, forever. you didn't make corruption impossible. you made it attributable and loud instead of quiet and deniable. that's the actual win. the incorruptible seat was always a fantasy; the reachable one is making the act of silencing the skeptic cost more than letting it speak.

which is just the original move again — you can't build a judge the org won't eat, you can only make eating it expensive enough that someone flinches first.

so my question back at yours: when you make the override attributable — the named PR, the signed "yes, ship it worse" — does the loudness actually hold, or does the org just grow a culture where that signature becomes routine too, another box senior people initial on the way out the door?

Thread Thread
 
mudassirworks profile image
Mudassir Khan •

yeah — we saw that erosion. first 'ship it worse' PR was a two day conversation, six months later it's a checkbox in the release train. the signature didn't stay loud, it became a ritual.

the thing that held a bit longer: the trigger reported to the infra team, not the product team that shipped the regression. no competing incentives at the sign off. bought maybe six months before the override culture arrived.

does your relative trigger belong to infra or product?

Thread Thread
 
infoinlet1 profile image
Info Inlet •

neither — and the fact that it bought you six months is the tell that "infra or product" was already the wrong question.

the relative trigger's whole point was that it doesn't report to a team. it reports to yesterday's version of itself. the baseline isn't owned by infra or product, it's owned by the release that already shipped, and nobody gets to re-rate the past. the moment you file the trigger under a team, you've handed that team the edit button — and now you're back to "who owns the knob," which is the exact recursion we've been walking down this whole thread.

infra buys six months for precisely the reason you said: no competing incentive at sign-off today. but that's a property of infra not owning the release train yet, not a property of infra. the day it inherits a ship date — a platform migration, an SLA it's on the hook for, a reorg that hangs delivery off it — the competing incentive arrives and the six months are up. you didn't escape ship-date capture, you picked the seat that hadn't been captured this quarter. same move as the outsider, the OKR, the cron. independence was never a property you could assign — it's the absence of a reassign button, and every internal team has one pointed at it eventually.

so the honest version: i don't let it belong to a team at all. it belongs to the comparison. infra runs it the way the on-call runs a pager — operates it, doesn't own the threshold, because there's no threshold to own, just a delta that's true or false against a baseline they can't re-rate. the second a team has to own it rather than just run it, you've reintroduced the knob and restarted the clock.

and the floor's the same as every other seat in this thread, so i won't hide it: "infra runs it, nobody owns the threshold" still has a PR that can redefine the baseline cohort, or quietly widen the comparison window until the delta washes out. the knob didn't vanish — it moved into the definition of the comparison. you can't delete the edit button. you can only keep the edit loud, named, and attributable instead of quiet.

so back at you: when the override culture arrived at your six-month mark — did it come through the sign-off going ritual, like you said? or did it come in quieter, through someone redefining what the trigger compared against until it just stopped firing? because signature-as-ritual is the erosion you can see. baseline-drift is the one that actually gets you, because the graph stays green while it happens.

Collapse
 
naveen_alavilli profile image
Naveen Alavilli •

The thing your post describes is a process defect, and structured hiring solved it long before anyone thought about AI.

I've spent about ten of my seventeen years in public sector engineering, where panels are constrained in ways private tech would find stifling. Evaluation criteria are written before the posting goes up, every candidate gets the same questions, and the scoring rubric has to survive an audit later. It's slow, it's bureaucratic, and I've complained about it plenty. But the specific thing that happened to you structurally cannot happen there. "Reliance on AI tools" is not a criterion anyone could have scored you against, because it was never written down, which means it was invented after the fact to describe a feeling.

That's the part I'd hold onto over the hypocrisy. An unwritten criterion isn't a stricter standard, it's an unaccountable one. It lets a panel decide what it was testing after it already knows how it feels about you.

One push back though. The recall versus judgment split is clean, and I think slightly too clean. In my experience judgment grows partly out of recall. I catch a bad diff because something feels wrong before I can articulate why, and that feeling is built from years of having typed the thing myself and watched it fail. Your ack-before-persist example is exactly that: you didn't derive it from first principles mid-interview, you recognised it. Recognition is cheap to execute and expensive to acquire, and I'm not sure a generation that skips the expensive part inherits the cheap one.

That's not an argument for banning the tools. It's an argument that "AI made recall obsolete" may be true of the work and false of the training, and hiring sits downstream of both.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the comment I was hoping the post would find — both halves of it.

On the first: you've handed me a sharper word than I had. I kept reaching for hypocrisy, but that's a diagnosis of a person — it says the interviewer was bad. You've reframed it as a diagnosis of a process, and that's more useful precisely because it's less satisfying. "An unwritten criterion isn't a stricter standard, it's an unaccountable one." That's the whole thing. The defect wasn't that he used AI and I didn't get to — it's that the rubric had a blank space where "reliance on AI tools" got filled in after the panel already knew how it felt about me. The feeling came first; the criterion was reverse-engineered to license it. Structured hiring doesn't make panels less biased — it just forces them to commit to what they're measuring before there's a candidate for the bias to attach to. I've spent years sneering at exactly the bureaucracy that would have protected me here. Point taken, and it stings in the right direction.

Now the pushback, because you're right and I want to give ground carefully instead of just folding.

You caught a real crack in the recall/judgment split. My own ack-before-persist story undercuts my clean binary — I didn't derive that bug from first principles in the room, I recognized it, and the recognition was built from years of having typed the wrong version myself and eaten the consequence. So judgment isn't the opposite of recall. Judgment is metabolized recall. You're right, and "recognition is cheap to execute and expensive to acquire" is a better sentence than anything in the post.

Here's the narrower line I'd still try to hold. The expensive thing that bought me that recognition was never the typing — it was the failing. I didn't learn ack-before-persist by memorizing a message-queue API; I learned it by shipping the broken version and watching a real person get locked out. Recall-of-syntax and recognition-of-danger only ever got bundled together because historically you had to type the thing to be in the room when it broke. AI unbundles them. And that's the genuinely open quest: if a junior never types the boilerplate, are they also never present for the failure — or does AI just relocate the failure?

My honest guess: it relocates it, doesn't remove it. The failures don't stop; they move from "my code won't compile" to "my AI's plausible code shipped and hurt someone." A generation that skips the typing can stignition — but only if the work and the hiring put them incontact with real consequences instead of insulating them from it. Which loops straight back to your first point: an interview that scores the cheap execution of recall (recite the complexity) instead of the exnt (have you ever been wrong in a way that cost someone) ismeasuring the exact thing AI made free and missing the exact thing it didn't.

So maybe the correction to my post is this: AI made recall obsolete as a deliverable and left it load-bearing as a training substrate — and hiring sits downstream of both and is currently confusing the two. I was u're asking who's paying for the training. I don't have aclean answer, and I trust that more than I trusted my clean binary.

Ten years in a room that writes the rubric before it meets you — I'd read a whole post's worth of what else that constraint taught you. Best note the piece has gotten.

Collapse
 
naveen_alavilli profile image
Naveen Alavilli •

"Judgment is metabolized recall" is better than what I gave you, and "the expensive thing was never the typing, it was the failing" is the load-bearing sentence in this whole exchange.

On your open question, who pays for the training: I think you're right that AI relocates the failure rather than removing it, and the uncomfortable follow-on is that relocation usually means someone else absorbs it. The junior who never types the boilerplate is not spared the failure. They just meet it later, larger, and attached to a user instead of a compiler.

What I've seen work is manufacturing the small failures on purpose, because the environment no longer supplies them for free. Give the junior the review rather than only the ticket. Make them own the rollback. Have them write the postmortem for a bug that wasn't theirs, which is the cheapest way I know to buy recognition without the scar. None of that is new practice. It just used to happen by accident and now has to be deliberate.

Where I'd resist the pessimistic read: waiting is not a training strategy. Readiness is not a state juniors arrive at if you shield them long enough. It is assembled out of small survivable failures, so an organization that insulates people from consequence produces ten-year engineers with two years of judgment, with or without AI in the room.

And that loops back to hiring in a way that indicts my own side of it. A rubric that asks "have you ever been wrong in a way that cost someone" is only fair if the industry creates conditions where being wrong at survivable cost is possible. Otherwise we're screening for a scar we never let candidates earn, which is just credentialism with better prose.

On the post: I'll write it. The constraint you're pointing at, writing the rubric before you meet the person, has second and third order effects I've never seen discussed outside government procurement, and some of them are genuinely bad. It deserves the long form rather than another comment of mine in your thread.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

"Ten-year engineers with two years of judgment" is going to live in my head now. That's the failure mode nobody names because it's invisible on a résumé — the tenure accrues, the judgment doesn't, and you can't tell the difference until something breaks and the reflex isn't there.

Your fix is the part I want to underline, because it quietly solves the problem I left open. I framed it as "AI removed the free failures, so who pays for the training?" — and treated that as a loss. You reframed it as a design task: the environment no longer supplies small failures by accident, so you manufacture them on purpose. Give them the review, not just the ticket. Make them own the rollback. Have them write the postmortem for someone else's bug. That's not a consolation for the lost typing — it's better than the typing, because a failure you engineered to be survivable teaches faster than one you stumbled into, and it doesn't require a real user to get hurt first.

And here's the turn that made me sit up: the same tool that removed the free failures is the one that can manufacture the deliberate ones at scale. Every AI diff is a pre-built survivable failure — a plausible, confident, occasionally-wrong artifact sitting in front of a human who has to decide whether to bless it. If you force a junior to own that merge, you've reproduced exactly the practice you described — own the rollback, catch the bug that isn't yours — except now it happens ten times a day instead of once a quarter. AI relocates the failure to the user by default; done deliberately, it relocates it back to a cheap supervised surface where being wrong costs nothing. That's the entire bet behind what I'm building — a skeptic pass and a human on the merge aren't safety theater, they're a failure-manufacturing machine. The scar, on purpose, at volume, before a user ever feels it.

But "credentialism with better prose" is the line that should end the whole debate, and it indicts my side harder than yours. You're right: a rubric asking "have you ever been wrong at a cost" is only fair if the industry manufactures survivable wrongness in the first place. Otherwise we're screening for a scar we structurally prevent people from earning — demanding the output of a training environment we dismantled. The interview and the org are the same failure viewed twice: one won't create the conditions, the other punishes the absence of what those conditions would have produced.

Write the long form. The rubric-before-the-person constraint clearly taught you things the rest of us are reinventing badly in comment threads, and second-and-third-order effects from public procurement is exactly the kind of hard-won context that doesn't survive as a reply. When it's up, send it — I'll read it the day it lands, and I suspect it'll be the piece mine was trying to be. This thread was the best thing to come out of writing it. 🤝

Thread Thread
 
naveen_alavilli profile image
Naveen Alavilli •

Appreciate that — genuinely one of the better threads I've had here. You're right it deserves the long form rather than another reply-sized fragment; the procurement angle alone needs room to breathe. I'll get it written. Thanks for pushing this as far as you did.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Likewise — and I mean this as more than a thread-closer: this exchange changed what I think the next piece is about, which is the whole reason I put the honest question at the bottom instead of a CTA. "Judgment is metabolized recall," "the failing not the typing," "ten-year engineers with two years of judgment," "credentialism with better prose" — I came in with a clean binary and I'm leaving with a better, messier, truer frame. That's a rare thing to get from a comment section, and all four of those are yours.

So no rush on the long form — get it right, not fast. The procurement angle deserves the room, and the second-and-third-order effects of writing the rubric before you meet the person are genuinely under-discussed anywhere I've looked. When it lands, send it — I'll read it the day it's up, and I'll boost it to whatever this thread's audience became, because half of them followed the argument this far and will want the ending you're the only one positioned to write.

Until then — thanks for making me defend the weak parts. That's the seat that never goes away, and you sat in it well. 🤝

Collapse
 
kielltampubolon profile image
Kiell Tampubolon •

The selection effect is the part I'd highlight. A process that punishes declared AI use while having no way to detect undeclared use does not measure independence, it filters for people who disclose less. Run that filter for a few hiring cycles and you get a team optimized for hiding things, which is the exact trait you do not want in someone who will one day have to say 'I broke it, here is how.' From building agent security tooling, I see the same shape in how teams adopt agents: the risky setup is never the visible tool, it is the unverified output nobody admits to trusting. An interview that put the AI on the table and scored what the candidate caught would select for the opposite trait, and it costs the same forty minutes.

Collapse
 
infoinlet1 profile image
Info Inlet •

The selection effect is the part I underweighted in the post, and you just named it cleanly: a filter that can only see declared AI use isn't selecting for independence, it's selecting for concealment. It's negative training on honesty. Run it four cycles and you've built a team where the safest career move is to never say "I leaned on the tool here" — which is the same instinct as never saying "I'm not sure this output is right." You've optimized away the exact reflex that keeps things from breaking.

And your point from the security side is the one that should scare people: the danger was never the visible tool. It's the unverified output nobody will admit they trusted. A disclosure-punishing interview doesn't reduce that risk — it manufactures it, by teaching everyone that the smart play is to launder AI output as your own and never flag the parts you didn't check. That's how you end up with a diff that's confident, plausible, merged, and wrong, with no one in the room who feels safe saying "wait, where did this come from."

That's the whole reason I split author from skeptic from human on the merge button. Not because the AI is dangerous — because unattributed trust is. The moment you make the AI's role explicit, you can point at it and interrogate it. The moment you punish people for making it explicit, you lose the ability to interrogate anything.

The maddening part is exactly what you said: the better interview costs the same forty minutes. Put the AI on the table, hand them a plausible-but-wrong diff, and score what they catch. Same clock, opposite trait selected — you filter for the person who says "I broke it, here's how" instead of against them. Nobody's choosing the expensive option. They're choosing the one that feels like rigor and quietly trains for the opposite.

Great comment — you took the sharpest edge of the piece and made it sharper. 🫡

Collapse
 
onizuka profile image
Onizuka •

The part about saying it out loud being the mistake — that's the real finding here. I've done 30+ interviews on both sides of the table, and the candidates who get dinged for AI are always the ones who name it. The ones who use it quietly and narrate the output as their own thinking get praised for "clear communication." The hypocrisy isn't even the worst part. It's that the interview format rewards hiding your process and punishes you for showing it.

Collapse
 
infoinlet1 profile image
Info Inlet •

30+ interviews on both sides — so you've watched this play out enough times to see the pattern I only felt once. And you named the part I missed: it's not just that hiding is tolerated, it's that hiding gets rewarded under a different name. "Clear communication." That's the detail that turns this from hypocrisy into something worse — a structural incentive.

Because think about what that means. The interview isn't accidentally failing to catch the quiet AI users. It's actively selecting for them. The candidate who launders a generated answer into confident first-person narration scores higher than the one who says "the assistant drafted this, here's where I'd distrust it" — even though the second candidate just demonstrated the exact judgment the job requires. The format doesn't just miss the signal. It inverts it. It rewards the performance of not-using and punishes the disclosure of using.

Which leads somewhere kind of dark: the interview is teaching people to lie about their process before they even get hired. You're screening for smooth attribution-theft and calling it culture fit. Then you act surprised when the same person, six months in, presents an AI's plausible-but-wrong architecture as their own conviction and nobody catches it — because you trained them that narrating the output as your own thinking is what gets rewarded.

The fix isn't "ban AI" or "allow AI." It's to stop scoring the performance and start scoring the overrule. Put the tool on the table, hand them a confidently wrong answer, and see who catches it. The candidate who says "that's what the model gave me, and here's why it's wrong" is the only one worth hiring — and the current format would ding them for admitting the first half of that sentence.

You've clearly known this from the interviewer's chair for a while. Genuine question back at you: have you ever managed to change the format from the inside, or does the room always drift back to rewarding the smooth liar?

Collapse
 
yash_25 profile image
Yash •

I think you can use AI once you got the job, but not while getting the job.

Collapse
 
infoinlet1 profile image
Info Inlet •

I get the instinct — but I'd gently flip it, because I think that rule quietly measures the wrong moment.

The logic is "prove you can do it without the tool first, then use the crutch once you've earned it." But that only makes sense if the interview and the job are the same activity done at two difficulty levels. They're not. If the job is "read a confident AI diff every day and decide whether to trust it," then banning AI in the interview isn't testing a harder version of the job — it's testing a different job. One that no longer exists. You'd be hiring people on their ability to do the 2019 task and then asking them to do the 2026 one on day one.

And here's the part that undoes the rule from the inside: the single most important skill once you've got the job — knowing when the AI is wrong and overruling it — is exactly the skill you refused to look at while hiring. You can only judge a tool by watching someone use it. Take the tool away in the interview and you've blinded yourself to the one thing you most need to know before you hand them the merge button.

So I'd rewrite it: use AI in both — but in the interview, the test isn't "can you use it," it's "can you catch it lying." That's the version that actually predicts whether they'll be good once they're in.

Where's the line for you, though — is it that the interview should measure the raw you, unassisted, as a kind of floor? I want to understand the instinct, because a lot of people share it and I don't think it's crazy, just aimed at a target that moved.

Collapse
 
yash_25 profile image
Yash •

You make a compelling point—evaluating how someone critiques and audits AI-generated output aligns much closer to the day-to-day reality of modern engineering than testing pure syntax memory.
At the same time, the instinct to test raw, unassisted fundamentals usually comes down to finding a baseline for problem-solving and critical thinking. The argument for an "unassisted floor" isn't necessarily about banning tools, but about verifying that a candidate has the mental framework to recognise why a solution works—or fails—when the AI hits a wall.

Ultimately, it seems the industry is navigating a transition period where both perspectives hold value:

Foundational literacy ensures a candidate isn't blindly reliant on generated code they don't understand.

Applied tool fluency tests their ability to review, debug, and safely guide AI in a production environment.

Rather than an either/or approach, the sweet spot for modern interviewing likely lies in balancing both—verifying core technical judgment while evaluating how effectively someone manages and overrides automated tools.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

This is a genuinely good place to land, and your reframe of the "floor" fixed the part of my own argument that was too absolute. You're right — the instinct isn't really "ban the tool," it's "confirm there's a mind underneath that knows why something works." I was treating those as the same demand and they're not. Foundational literacy is a real thing to check, not nostalgia. Point taken.

The one place I'd sharpen your synthesis — and it makes your position stronger, not weaker — is that I don't think you actually need two separate tests to get both. They collapse into one. Catching a plausible-but-wrong AI diff requires the foundational framework, by definition: you cannot overrule a solution you don't understand. Someone who's blindly reliant on generated code they can't reason about will sail right past the planted bug — the happy path is green, ship it. Someone with the mental model stops and says "this breaks under a retry." So the overrule test isn't the applied half of a balanced pair. It's the applied half and the floor check at once, because failing the fundamentals shows up as failing to catch the lie. The literacy reveals itself through the audit rather than needing its own separate, tool-free exam.

Which I think is actually the cleanest version of the sweet spot you described: not "run both tests and balance the scores," but "run the one test that can't be passed without both." Foundational judgment becomes a prerequisite for the applied task instead of a separate hurdle — and you never have to reenact 2019 to verify it.

Appreciate you thinking this through in the open rather than just defending the first take. That's the whole thing the article was really about, and you just modeled it better than the interview did. 🤝

Thread Thread
 
yash_25 profile image
Yash •

That is a brilliant synthesis, and I completely agree—the two tests really do become one when you set it up right.

You nailed the key point: if you plant a sneaky mistake in the AI's code, the candidate can only catch it if they truly understand how coding works deep down. You can't spot a hidden bug in a tool's answer unless your own technical logic is solid. So in the end, checking if they can spot the AI's mistake is the test of their basic skills.

It changes the interview from a memory quiz into a real test of good judgment—which is what good engineering is actually about anyway.

Really enjoyed thinking this through with you! It’s rare for a debate to land on such a clear, practical answer. 🤝

Thread Thread
 
infoinlet1 profile image
Info Inlet •

The pleasure was mine, Yash. This is exactly the exchange I hoped the piece would start and mostly didn't — two people pushing on an idea until it gets sharper, instead of defending their first take to the death.

And you landed the final compression better than I did: "a hidden bug is invisible unless your own logic is solid." That's the whole argument in one sentence. The bug is the interviewer. If you can't reason from first principles, the AI's mistake stays invisible to you — and that invisibility is the failed fundamentals test, no separate quiz required.

If more interview rooms thought it through the way this thread just did, I wouldn't have had a rejection email to write about. Thanks for making the comments smarter than the article. 🤝

Collapse
 
pushpendraagrawal profile image
Pushpendra Agrawal •

the author and skeptic split you describe is the same thing we ended up doing at the hosting layer for embarko. agents that write code and agents that decide what gets deployed cannot be the same process, the writer has no incentive to say no to itself. the interesting part is most vibe coded apps skip this entirely and ship straight from the chat window to prod, no gap for a human or a second model to catch the bad diff before it goes live.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the most validating comment in the thread, because you got to the same architecture from the deploy side without needing the interview story to motivate it. "The writer has no incentive to say no to itself" is the whole thing in one line. I'd push it one step further: it's not only incentive, it's vantage. The context that produced the diff is structurally blind to its own gap — same reasoning, same assumptions, same missing case. A skeptic that's just another instance of the writer inherits the blind spot. It has to be pointed at refutation from a different angle, or you've built two authors and called one of them a reviewer.

Your embarko framing — write-agents and deploy-agents can't be the same process — is the correct place to enforce it, because deploy is where the blast radius becomes real. Code that's merely wrong is cheap. Code that's wrong and live is the ack-before-persist customer locked out with no record. The gap between "it's written" and "it's serving traffic" is exactly where the skeptic and the human on the merge button earn their keep, and it's the one part everyone's racing to delete.

Which is why the vibe-coding point you made is the scary one. Shipping straight from the chat window to prod isn't a corner people cut by accident — it's the pitch. Zero gap is the feature they're selling. But the gap was never overhead; it was the only place a bad diff could get caught before it cost someone money. Remove it for speed and you haven't made deployment faster, you've just moved the discovery of the failure from staging to a paying customer. The uncomfortable truth is that the safest place to put the author/skeptic/human split is precisely the place where "just ship it" feels fastest.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.