DEV Community

Tahir Almas
Tahir Almas

Posted on • Originally published at ictlms.net

AI Training Went Everywhere. Proof Did Not. What a Smart Online Exam Measures

Originally published at ictlms.net

Most organizations can tell you exactly how many people completed AI training last quarter. Very few can tell you which of those people can now do the work. A smart online exam is what closes that distance, and the numbers coming out of 2026 suggest almost nobody has built one yet.

The number that should be bothering your board

IDC expects more than 90% of global enterprises to hit critical AI skills shortages by 2026, and puts the price of sustained skills gaps at up to $5.5 trillion in delayed products, quality problems, missed revenue and lost ground to competitors. That is not a training-department number. That is a business number.

Sitting next to it is a stranger one. Around 94% of CEOs and HR chiefs name AI as their top in-demand skill, yet only about a third of leaders believe they have actually prepared their people. Only a third of organizations describe themselves as fully ready to adopt AI-driven ways of working, and only a third of employees report receiving any AI training at all in the past year.

So the spending is happening and the confidence is not arriving with it. When you line the survey figures up, the shape of the problem gets obvious.
Four percentages, one story. The fall from 82 to 18 happens entirely in the space where nobody is testing anything.
Roughly 82% of enterprise leaders say their organization provides AI training and 68% report a dedicated program. Ask the employees who sat through it and only about 18% say it prepared them to work independently. Companies are spending in the region of $1,200 a head per year on this.

Completion is not capability, and everyone quietly knows it

Here is the uncomfortable bit. The metric almost every organization reports upward is completion. Seats filled, modules finished, quiz passed, badge issued. It is easy to collect, it goes up and to the right, and it measures attendance.
Same person, same course. One column answers whether they turned up, the other answers whether they can do the job.
A ten question quiz at the end of a module tests whether someone remembers the slide from eleven minutes ago. That is recall, and recall decays fast. Six weeks later most of it is gone, and the completion record still says complete, forever.

My honest view is that this is not laziness. Completion is measured because it is cheap, and capability is not measured because building a real assessment used to mean writing scenarios, marking them by hand, and finding humans with time to do it. That constraint has changed, which is exactly why the excuse has run out.

What a smart online exam has to test if it is going to mean anything

Testing AI skills with multiple choice is close to useless. The whole point of the skill is judgment under messy conditions, and you cannot get at judgment with four options and one right answer.

The tests that predict real performance tend to do four things. They hand over a task rather than a question, using an actual document, dataset or customer message from the business. They let the candidate use AI tools during the exam, because banning the tool you are assessing them on makes no sense. They plant something wrong in the AI output and see whether it gets caught, since spotting a confident wrong answer is the skill that separates competent from dangerous. And they ask for the reasoning, not just the result.

That last one is where AI-assisted grading earns its place. Marking a few thousand written justifications by hand is what killed this idea in the past. A model can read them all, group the reasoning patterns, flag the weak ones, and hand a human a short pile to review instead of the whole stack.

Keep the proportions sane

There is a real risk of overcorrecting here. An internal readiness check does not need the machinery of a licensing exam. If the result feeds a training plan, keep it light: open book, real tasks, no camera, no lockdown browser. Save the heavier controls for the cases where the certificate travels outside the company and carries weight with a customer or a regulator.

We have argued before that AI should flag and humans should decide, and that applies just as much to grading as it does to integrity monitoring. An automated score that nobody can explain to the person who received it will not survive its first appeal.

Where this fits with the systems you already run

Nobody needs to replace their LMS to do this, and I would push back on any vendor who suggests it. Course delivery, enrolment and records are fine where they are. What is usually missing is the assessment layer that produces evidence.

ICTExam is built to sit in that gap. It connects into an existing platform over LTI 1.3, so learners launch an exam from the course they are already in and results flow back without anyone exporting spreadsheets. If you want the full picture of what it does, the feature list and the integrations page cover it in more detail than a blog post should.

A 60 day version you can actually run

  • Pick one role, not the whole company. Customer support, or analysts, or the sales team. A single role gives you a clean signal and a short argument.
  • Write down four things that role should be able to do with AI. If you cannot name four, the training probably had no target either.
  • Build three tasks from real work. Last month's actual tickets or documents, lightly anonymized. Invented scenarios produce invented results.
  • Salt one task with a plausible AI error. A wrong figure, an invented policy, a citation that does not exist. This single item will tell you more than the other two combined.
  • Run it open book with the tools allowed and give people a time box rather than a lockdown.
  • Grade with AI, review a sample by hand. Read every flagged script and a random 10% of the rest, then check whether the machine and the humans agreed.
  • Retest the same group at six months. The retest is where you find out whether anything stuck, and it is the step everyone skips.

Run that once and you will have something no completion report can give you: a defensible statement about what a specific group of people can do, with the evidence attached.

Frequently asked questions

Should employees be allowed to use AI during an AI skills exam?

Yes, in almost every case. You are assessing how well they work with the tool, and taking it away tests something else entirely. The exception is a foundational test where you need to know they understand the underlying material without help.

Is AI grading reliable enough for a workplace assessment?

For structured, scenario-based answers with a clear rubric, it is good enough to do the first pass and it is dramatically more consistent than a tired human at 6pm. Keep a person reviewing flagged and borderline scripts. Fully automated grading with no human in the loop is not something we would recommend for anything that affects someone's job.

Do we need proctoring for internal assessments?

Usually not. If the result guides training rather than gating pay or promotion, cheating mostly harms the cheat and pollutes your data. Add controls when the stakes rise, and tell people clearly what is being monitored when you do.

How often should people be retested?

Every six months works well for AI skills, mainly because the tools change so fast that a test from a year ago is measuring a product that no longer behaves the same way. Annual is the minimum I would defend.

Can this run alongside Moodle or another LMS?

Yes. LTI 1.3 is the standard route and it means learners never leave the platform they know. See the LTI integration guide for how the launch and grade passback work.

What size organization is this worth doing for?

Below roughly 50 people a manager can usually judge capability by watching. Past that, informal judgment stops scaling and starts being wrong in ways nobody notices.

Related resources

If you already run the training and just need the proof, the quickest way to see whether this fits is to put one real task in front of it. Take a look at the ICTExam demo, or start with the platform overview if you are still scoping the problem.

Top comments (0)